Showing posts with label error-proofing. Show all posts
Showing posts with label error-proofing. Show all posts

Thursday, August 14, 2025

Poka-yoke, or error-proofing

A few days ago, a friend of mine had a minor surgery. (He's recovering just fine.) One of the interesting parts of the preparation, though, was that the surgical team asked him to mark on himself where they were supposed to operate.

Wait, what? Isn't that something they should already know before they get this far?

This is a stock photo and not my friend.
But you get the idea.
 
Well yes, sure. But think about it. One of the worst possible mistakes on the surgical table—short of losing the patient altogether, I mean—is the risk of operating on the wrong part of the body. And while it's very rare, nonetheless it has happened: you can google stories of people who went in to have their left leg operated on, and the surgeon operated on the right leg instead. The easiest way to prevent that kind of mistake is to ask the patient—who presumably knows where it hurts—to take a big blue marker and put a circle or an arrow on the spot that needs attention. It's easy to do, and it's just one more line of defense against error.

In fact, error-proofing (sometimes described using the Japanese term poka yoke, or ポカヨケ) is a fundamental Quality method. If you can make an error impossible (or very unlikely) then you save all the time and expense of having to fix it later. And of course in the case of a surgical error, "time and expense" doesn't begin to cover it. 

So how do you error-proof a process? There are almost as many answers as there are processes. But consider the process of cutting wooden rods or metal poles to a fixed length. In principle you could ask the cutter to measure each rod one at a time, and then cut it. This approach is slow, and it also introduces a lot of chance for error: every time the cutter measures a new pole, the measurement is plus-or-minus a certain amount. At the end of the shift, there's a good chance that no two poles are the same length. But you can error-proof the operation by giving your cutter a jig or fixture to seat the poles in. Then if he cuts across the top of the fixture, the poles will all be the same length.

You can go one step farther. Use a table saw with a rip fence. Then you can align every pole exactly the same. Depending on the exact nature of what you are cutting, you may even be able to cut several at once.

By extension, you can think of a lot of safety-guards on machinery as error-proofing devices. What else is a hand-guard, after all, than a device to prevent you from committing the error of reaching your hand inside the equipment while the gears are turning? The simplest hand-guard does exactly that. A more elaborate form is designed so that the machine cannot operate unless the guard is in place: this prevents you not only from reaching your hand into the gears, but from removing the hand-guard and trying to operate the machine without it. And overriding or defeating a machine's safety equipment is another kind of error.

One of the earliest works
describing double-entry
bookkeeping
In an entirely different vein, consider the invention of double-entry bookkeeping in the late Middle Ages. Double-entry bookkeeping requires that every transaction has to be entered into two different accounts. 

For example, if a business takes out a bank loan for $10,000, recording the transaction in the bank's books would require a DEBIT of $10,000 to an asset account called "Loan Receivable", as well as a CREDIT of $10,000 to an asset account called "Cash".*

Strictly speaking double-entry bookkeeping is an error-detection tool rather than an error-prevention tool, because the sum of all credits to all accounts must always equal the sum of all debits from all accounts. If these two sums ever fail to match, there is an error somewhere. But the difference between error-detection and error-prevention is more terminological than real. If you detect an error soon enough, you correct it before it has time to compound into larger errors farther along in the process. 

In cases where the cost of error is smaller, the methods for error-proofing are correspondingly less formal. Another friend works in retail. Sometimes customers call in saying, "I was there yesterday looking at the products against the far wall, and I want to buy the one on the left. Is it still there? Can you set it aside for me, and I'll pick it up later?" Of course nobody knows which one was "on the left" back when the customer was in the store, because products might have been rearranged since then. But it's easy enough for a clerk to walk over to the far wall, take a picture, and text it to the customer. "Is this the one you want? Great. It'll be waiting by the register when you come in." The potential risk in case of error is not as great as in surgery, but it's a simple way to double-check and it makes the customer happy.

And that, of course, is the point.

__________

* Quoted from Wikipedia, "Double-entry bookkeeping."  

      

Thursday, April 14, 2022

"Sage powers" and learning to listen

Last week I attended another great webinar. It was hosted by Angie Alexander, of Corporate Catalyst Consulting, and she was interviewing Jeff Griffiths of WorkForce Strategies International. Regular readers will remember that I wrote about Griffiths last fall: in this post where I described his webinar "People Before Process," and then in follow-on posts (here and here) where I picked up a later question that he and I discussed afterwards and worked my way through it.

This webinar was called "Getting Sh*t Done! – Optimizing Performance through Human Centric Sage Leadership."* Ms. Alexander's starting point is the recognition that our internal attitude towards our work has a huge impact on our outcomes. Maybe that doesn't sound controversial when I state it so broadly, since the general message dates at least to the time of Epictetus. But Alexander then developed the point in some detail by identifying ten "saboteurs" (ideas or attitudes that get in the way of our working well) and five "sage powers" that combat the saboteurs and allow us to work with calm, clarity, and happiness. These five sage powers are:

  • empathy, an unconditional love of yourself and others
  • exploration, based on a love of learning
  • innovation, a creative willingness to think of all the possibilities
  • navigation, a search for purpose and meaning that asks "What's really important?"
  • activation, moving forward without all the distractions
During the subsequent conversation, Griffiths picked up each of these themes in turn and gave examples from his own career. To take just the very first one, he explained that without empathy it's easy to react to a difficult situation by feeling you are surrounded by idiots. But if you slow down, take a deep breath, and accept them unconditionally, then you can start to see where these people are coming from and they no longer look like idiots. Then you can work with them in a productive way. 

For my part, I spent the whole webinar thinking about the Quality business. I saw at least two places where the crossover is very strong.

One of these is in the area of problem-solving. You remember back in January I discussed the principle that "There is no such thing as human error." At the time I called this a "motivational slogan," and I emphasized how it is a useful approach if you want the cooperation (which you need) of people who were on the scene when an accident occurred. But this approach is just empathy in action, a willingness to see the events through the eyes of another person as a first step towards error-proofing the operation for the future. And after all, if your problem-solving exercise comes up with a solution that would work for you (if only you were the one on the scene) but that won't work for the people who actually do the job – in other words, if your solution lacks empathy for the people who have to implement it – then you really haven't solved anything and you can count on the problem to recur.

The other is in the area of auditing. One of the most dangerous temptations for an auditor is the desire to start telling your auditees how to do everything differently. In your mind it's simple: you know the management system standards, you've seen how those standards are implemented by hundreds of other companies, and these folks are simply doing it all wrong! And from the other side, pretty much every organization that goes through regular external audits has met one or more auditors who give into this temptation. I've talked before about how the line between auditing and consulting gets particularly fuzzy during internal audits; but during external audits it is important to keep it clear. And even in internal audits, you can't abuse your authority as an auditor just because the manager of this department stubbornly insists on doing things in a way that looks crazy to you.

For this reason, I have found it very useful when conducting audits to slow down, keep quiet, and listen. I'll ask how this or that part of the system works, and then just let the auditees talk. At first glance I might think they are doing it all wrong; I tell myself, "If I had designed their system, I would have done it differently." But it's their system, not mine. So if I wait and listen, after a while I can start to see the system through their eyes. I can start to see how the strange bits all hang together as a system. And I can see that my first objections were often based on my own failure to understand. At this point I may have some thoughts about how certain aspects can be improved – "Is there a risk because you aren't keeping minutes from this or that regular meeting?" – but I'm no longer going to ask them to redesign the whole thing from the beginning.

At its best, then, auditing deploys – or should deploy – all five of Alexander's "sage powers": empathy, to understand the client's system on its own terms; exploration and innovation, to look past the auditor's own preconceptions and see what the client is trying to do; navigation, to understand which topics really can put the client's system at risk; and activation, to proceed from there and find the nonconforming gaps (if any) which it will really add value for the client to fix. Of course if you find other nonconformities along the way you have to report them. But you provide the best service if you focus where it hurts.

It's not always easy, and at the end your list of findings may be shorter than it would have been doing it in a more directive way. But it's more helpful to work this way, which is the whole point. Besides, we're not paid by the finding.  😀        

__________

* Yes, the second word in the title really did have an asterisk instead of the vowel.     

Thursday, February 3, 2022

What about human error? Part 2 of 2

Last week I talked about the concept of "human error" — and especially about why we in the Quality business always insist that "There's no such thing as human error" when it's obvious that there is. I argued that while of course we all know that humans make mistakes, that's never the place to stop in an incident-investigation: focusing hard on human error makes your participants clam up if they are afraid of being blamed, and it shuts off the chance to find systemic improvements that could make future mistakes less likely. Another way to say this is to say that human error is a symptom but not a cause. Somewhere in your system, there is something else that triggered or allowed the human error to happen, and that's the thing you want to find and control.

But if human error is a symptom then we really need to understand what kinds of human error we might encounter, because each different kind of error is probably a symptom of a different cause and therefore has to be treated in a different way. If you go to the doctor because of pain, he'll treat you differently depending whether the pain is in your head or your elbow.

Fortunately this work has already been analyzed and tabulated. The information that I provide below comes from a website page owned and administered by the United Kingdom's Crown Health and Safety Executive, and I gratefully acknowledge permission to use it, as follows: This blog post contains public sector information published by the Health and Safety Executive and licensed under the Open Government License. The full text of the Open Government License for public sector information can be found here.

So, what are the different kinds of human error?

In the first place, a failure is either:

  • Inadvertent (an error)
  • Deliberate (a violation)
Errors can be either:

  • Action errors (where our action is not as planned)
  • Thinking errors (where our action is as planned but we planned the wrong thing)
Action errors can be either:

  • Slips (where we do something wrong)
    • Example: Flip a switch up instead of down.
    • Example: Transpose digits during data entry.
  • Lapses (where we fail to do something right)
    • Example: Forget to signal before turning at an intersection.
    • Example: Skip a step in a safety-critical procedure.
Thinking errors can be either:

  • Rule-based mistakes (where we misapply a good rule or else apply a bad rule)
    • Example: Misjudge passing the car in front of you, because you would have plenty of room if you were in your own car, but you are in your friend's car which has a lot less power.
  • Knowledge-based mistakes (where we have no rules and try to figure it out from scratch)
    • Example: Rely on an out-of-date map to plan your route through an unfamiliar town.
Finally, violations can be either:

  • Routine (where there is no meaningful enforcement, so everyone ignores the rule)
    • Example: A lot of cars on the freeway ignore the posted speed limit.
  • Situational (where we cut corners in certain cases, because of features specific to those cases)
    • Example: You document your design reviews scrupulously for new designs, but never document them for subsequent Engineering Changes because they always look so obvious you can't see taking the time to write them down.
  • Exceptional (where we take a calculated risk in breaking the rules in order to address a highly-unusual situation)
    • Example: A huge production order in your factory is due by Friday. On Tuesday, one of your machines comes due for preventive maintenance, but you keep it running (rather than shutting it down the way you are supposed to) because that's the only way to meet the production deadline.
That's the list.

Now what do we do about them?

Notice first of all that an approach which works for one kind of failure really won't help with another kind. Requiring someone to fill out a checklist is very likely to prevent lapses, but it will be completely useless in preventing (say) exceptional violations. A checklist, after all, helps the memory, and in case of a lapse the operator just forgot a step. But in the case of an exceptional violation, the operator knows perfectly well what steps he is choosing to skip — and why — so a reminder (in the form of a checklist) will make no difference at all.

The kinds of steps to consider are things like these.

You can download a useful reminder table, which summarizes all this information and more, from my Downloads Page.

     

Thursday, January 27, 2022

What about human error? Part 1 of 2

It's a commonplace in the Quality business that any time we start a problem investigation, we insist that "there is no such thing as human error." I say it in this post here. But what does that mean, anyway? And is it true?

At a superficial level, at any rate, it looks like there is something wrong. About a month ago the topic came up in this post and this one on LinkedIn. If the links don't open for you, the basic point is made by Christopher Paris, who points out that obviously all errors are made by humans! After all, they certainly aren't made by space aliens.

Obviously Paris is right that errors are made by human beings and not space aliens. But sometimes I think he's a little too hard on those of us in the Quality industry.* For myself, I've always taken the principle about human error as a motivational slogan rather than a statement of fact, and I think that in a pragmatic sense it performs two roles.

First, if you want to do a decent root-cause analysis, you have to get all the facts. This means getting the cooperation of whoever was there on-site when the problem happened. Now if this employee thinks you're going to blame the whole thing on him and his errors, he's not going to tell you a thing. So you start off by saying that the problem has to be with the system, not with him, and you just need his help to figure out how to improve the system. With luck this will put him at his ease, so you can make progress.

Second, sometimes when your problem-solving team is in the middle of its work, you'll have someone who really wants to get back to his desk to work on something else instead. So he says, "Look, this whole accident was caused by human error. Next time we just have to try harder, that's all. So can we wrap it up and get back to our real jobs?" The problem is that "trying harder" has never solved anything. Often — nearly always, in fact — there is something that can be improved in the system to make it easier to do the job right and harder to make a mistake. So to keep your team from giving up too early, you remind them that "there is no such thing as human error," and if they really believe in "trying harder" then the problem-solving team should try harder to find a systemic cause.

What do I mean by a "systemic cause"? It's the kind of thing I talked about here (and then expanded on in the next two posts here and especially here). If someone made a mistake out of ignorance, see if you can improve your training system. If someone made a mistake because his hand slipped, see if you can get him a tool that makes the work easier. If someone forgot that those drums were filled with nasty waste until he almost dropped his lunch in one of them (let's say it was a "near-miss" and nothing bad actually happened to the lunch), see if you can label the drums or put up signs. Those are all system-level improvements.

At the same time, it's important to notice something else. You remember that there's no such thing as a perfect process, and in the same way there's no such thing as a perfect system. There's even an old joke that says, "You cannot make things foolproof because fools are so ingenious." So while the problem-solving team always has to look for additional system improvements, the organizational management has to emphasize improving the overall competence of every employee. This is because, as we discussed a couple months ago, good people can work under bad processes a lot better than bad people can work under good processes. So the best way to error-proof your operation, so far as you can, is to strengthen both.

Next week we'll look at a typology of human errors, and at the preventive measures which work for each one. It turns out there are several different kinds of human error, and the measures which prevent this kind are no help at preventing that kind. Join me.     

__________

* In fairness, his stated purpose is to motivate us to pull up our socks on a number of basic issues, so it's to be expected that he not go easy on us.

      

Thursday, January 6, 2022

Finding root causes, Part 2: Reaching farther

When something goes wrong and you are investigating what corrective action to take, how many root causes do you have to find? Typically there is more than one. We've talked over the last couple of weeks about what a root cause is and how to find it, and it's not unusual for the logic path of a 5-Why to branch. For example: The fire started because there were oily rags and at the same time there was also a spark — either one of them alone would not have done it. So now you have to ask "Why were there oily rags?" and also (on a separate branch) "Why was there a spark?". 

But even in simpler cases you may need to follow several different paths if you want to get a full picture of what's going on. To understand why, remember that a Quality system is all about getting what you want, and that means minimizing the extent to which you are derailed by problems. You can do this in three ways, and a good Quality system uses all three: 

  1. When a problem occurs, fix it.
  2. Looking downstream from the problem (if it has already occurred), catch it.
  3. Looking upstream from the problem (if it hasn't occurred yet), prevent it.
And so a really thorough investigation of the root cause of some problem takes all three perspectives into account. 

  1. You want to know what happened and why, so you can fix it and make sure it never happens again. This is what we have been talking about up till now.
  2. You want to know how to catch it, which means asking a second question.
  3. And you want to know how it was possible in the first place, which means looking for a second kind of answer.
In the rest of this post, I will walk through both of these enhancements.

One caution, before I begin: don't go crazy with any of this. Remember that your Quality methods have to be pragmatic: they have to serve you, and not vice versa. Add these enhancements so far as they are useful — and often they truly are very useful — but keep your level of effort proportional to the problem you are trying to solve.

Two questions

One way to make your investigation reach farther is to ask two questions instead of one, and to do a 5-Why analysis on each. The two questions are, "Why did the problem happen?" (which we have already discussed) and also "Why didn't we catch it in time?" 

The easy example is to think of a machine producing widgets that gets out of alignment, so that we start shipping crooked widgets to our customers. 

  • The first question asks about the machine: how did it get out of alignment? 
  • But the second question asks why nobody caught the problem in time: why didn't the inspectors at the end of the line see that the widgets were crooked and send up an alarm?

The point is that a working Quality system is built on the premise that things go wrong: machines break down, people make mistakes, and so on. As I noted above, a Quality system is designed to prevent mistakes before they can happen, and also to catch mistakes after they do. So if you've got a Quality system in place and a crooked widget slipped through anyway, there must have been several points of failure. Otherwise the problem would have been caught and corrected in the normal course of the workday.

If you ask about both the occurrence of a problem and its non-detection, that's sometimes called a "2 x 5-Why." And of course it makes your overall Quality system more robust and resilient, because it helps you to catch problems better as well as just fixing and preventing them. It enhances your investigation by adding the downstream perspective.

Two kinds of answers

The other way to reach farther is to look for two different kinds of root cause and not just one. The two kinds are "technical root cause" (which is the kind of root cause we have discussed up till now) and "managerial root cause." The idea behind this second one was summarized for me once by a senior colleague at one of the places I used to work — he was from another division, but I was lucky enough to get personalized training time with him — who said: "If you look at it right, everything that ever goes wrong in any plant is the fault of senior management."

Wait, what? How is that possible? I tried to make up some examples to prove him wrong.

  • What if some employee doesn't know what he's doing? Then the training process has broken down. And senior management set up the training program — either that, or they hired the person who did. So either they set up a faulty program, or they hired the wrong person.
  • What if a piece of equipment breaks? That equipment should have been covered by a preventive maintenance program: somebody should have been assigned to go around at regular intervals to check how the equipment is holding up, calibrate it if necessary, and then clean it and oil it. If that program had been in place, the employee responsible for maintenance would have seen that the part was getting worn and ordered a new one. But he didn't do it ... because there was no program ... because senior management either failed to set it up or failed to tell the Operations Manager to set it up.
  • What if an employee is measuring where to cut and the ruler slips? Isn't that a case of "Accidents happen"? Not at all. Why does he have to measure the cut with a ruler, when everybody knows that rulers can slip? The cutting operation should have been error-proofed by giving him a fixture to use: shove the thing to be cut until it is snug against the fixture, and then cut along the edge. He doesn't have to measure anything, and he gets the right length every time. But nobody built a fixture to error-proof the job, because — again — senior management didn't make sure that it happened and didn't hire an Area Lead who knew about these things. 
It went on like that for a while. Finally I got the point.  

Up above, the issue was that not only did something go wrong, but the system failed to detect it afterwards. In this case, the issue is that not only did something go wrong, but the system had to allow it to go wrong. That means that someone failed to set up the system correctly, or failed to execute some system-level task with an appropriate level of diligence. Either way, that's a managerial responsibility.

Think again of the widget machine we discussed above, the one that is out of alignment and making crooked widgets. 

  • The technical root cause might be that a certain part wore out, or perhaps the machine's design could be improved in such a way that it is less likely to slide out of alignment in the future. Either of these could be a valid cause and something we want to fix. 
  • But the managerial root cause has to do with flaws or gaps in how the system was set up by human beings. If a part wore out, the machine should have been serviced under a preventive maintenance program that would have found the part and replaced it in time. If the design was faulty, the development process which designed the machine in the first place should have foreseen that flaw (maybe by using an FMEA) and used a better design from the beginning.

In brief, this enhances your investigation by adding the upstream perspective.

As an aside: In non-industrial applications, it is not always obvious how to apply both of these elaborations, but it is usually worthwhile to think about it. And if you see a place where it can help, then use it.


I said before that if you ask about both occurrence and non-detection, the method is sometimes called "2 x 5 Why." If you do that and then also ask for both technical and managerial root causes, it is sometimes called "2 x 2 x 5 Why." And sure enough, in this case you really are asking four different questions, and you are looking for some meaningful and actionable answer for each one:

  • What is the technical root cause why the problem occurred?
  • What is the managerial root cause why the problem occurred?
  • What is the technical root cause why we didn't detect the problem in time?
  • What is the managerial root cause why we didn't detect the problem in time?

But the name isn't the important thing. The important thing is to understand what really caused the problem, so you can fix it. And if you can give good, actionable answers to all four questions, you have a really good framework to make sure this problem — and anything like it — never happens again.

     

Five laws of administration

It's the last week of the year, so let's end on a light note. Here are five general principles that I've picked up from working ...