Most five whys sessions stop at two or three. The name promises five for a reason: the first couple of answers are almost always symptoms, and the useful cause is usually further back than whoever's asking wants to go.
A real chain, not a toy one
A stamping line starts producing a spike of cracked brackets.
Why did the bracket crack during forming? The die's radius had worn down enough to concentrate strain at the bend instead of distributing it.
Why did the worn radius go uncaught? The preventive maintenance schedule checks die dimensions every 250,000 hits, but this die had gone back into production after an unrelated tooling swap, past the interval, without being remeasured.
Why did it go back into production without a check? The changeover procedure only requires a dimensional check when a die is pulled for wear-related reasons. A swap for an unrelated job isn't classified as a wear-pull, so no check was triggered.
Why does the procedure draw that distinction? The check adds about 15 minutes, and that time isn't budgeted into the changeover target for a routine swap. Wear-pulls get the time because they're expected to need it; swap-pulls don't.
Why is the changeover target tighter than the time a safe reinstall actually takes? The target was set from the press's OEE goal, without input from tooling engineering on what a dimensional check requires.
The fifth why doesn't land on "the die was worn" or "the operator missed it." It lands on a target-setting process that excluded the people who knew what a safe reinstall costs in minutes. That's a fixable root cause. "The die was worn" is not, it's a symptom that will recur under a different die number the next time a swap happens under time pressure.
Where the method breaks
Five whys produces a single chain, and manufacturing failures are rarely single-cause. The cracked bracket above could plausibly have had a second contributing thread, material lot hardness, or press tonnage drift, running in parallel with the die-wear thread. A linear five whys interrogation follows one branch to its end and stops; it doesn't naturally surface that a second, independent cause was also active. If the real failure has two or three contributing factors, a five whys session that only asks about one will produce a real, well-reasoned, and incomplete answer.
It also stops at the first plausible answer, not the correct one. Each "why" invites whoever's in the room to supply an explanation that sounds right, and once it sounds right, the session moves to the next question instead of verifying it against data. A worn die radius is a plausible answer to a cracked bracket. It's also the kind of answer that's easy to accept without pulling the actual dimensional log, and a five whys chain built on an unverified link at step one produces four more steps of reasoning about the wrong cause.
That's why the method is only as good as the person running it. Getting to a target-setting process instead of stopping at "operator error" requires someone in the room who knows that changeover time budgets get set by OEE targets, and who's willing to keep asking after the second answer feels satisfying. A facilitator without that depth stops at step two and calls it done.
When to escalate
Reach for a fishbone (Ishikawa) diagram when you suspect the failure has multiple contributing categories, machine, material, method, operator, measurement, environment, rather than one linear thread. A fishbone doesn't replace the questioning, it forces you to check each category instead of following only the first one that comes to mind. Reach for a fault tree when the failure is safety-critical or has multiple independent paths that could each trigger it, and you need Boolean logic (this AND that, or this OR that) to map which combinations actually cause the event, not just one chain that happened to be true this time.
Where AI actually helps, and where it doesn't
Pattern detection across many failures finds correlations a single five whys session, run on one bracket crack, structurally cannot: a die wearing faster on second shift, a supplier lot correlating with a defect three operations downstream, a failure mode that only shows up when two marginal conditions coincide. Dana used ML-driven root cause analysis from Acerta to cut axle rework 65%, landing at a 4% resulting rework rate and an estimated $2.5-3M in savings, by finding correlations across a much larger failure population than any single investigation session could hold in view. Owlet cut the time to validate a root cause 8x, saving $953K a year, by using the same kind of pattern detection to test hypotheses faster.
Neither replaces the judgment. A model can tell you that a correlation exists; it can't tell you that the correlation is causal, or that fixing it is worth the disruption, or that the real fix is a target-setting process nobody wrote down. That decision stayed human in both of the deployments above. What changed was how much evidence was available before the human had to make it.
Where to go next
- The quality control use cases and benchmarks pages cover outcomes across the full deployment corpus these root cause tools feed into.
- Vendors building pattern-detection and inspection software are ranked on the quality inspection software hub.
- Zaleco moved quality control in-process and cut scrap 20%, catching problems at the source rather than investigating them after the fact.