RCFA for Manufacturing: 5 Why & Fishbone Analysis

By Johnson on August 26, 2026

root-cause-failure-analysis-manufacturing-5-why

A conveyor belt snaps at two in the morning and takes the entire line down with it. The crew swaps the belt, restarts production, and everyone moves on. Six weeks later the same belt snaps again, in the same spot, on the same shift, and the instinct is still to swap the part and restart. Nobody ever asked why the belt failed the first time, because the pressure to get the line running again always wins against the fifteen minutes it takes to ask. Root cause failure analysis exists specifically to break that cycle, and choosing the right method for the failure in front of you is most of the battle, something our reliability specialists walk plants through every week.

Root Cause Failure Analysis

Stop Replacing the Same Part Every Few Weeks

5 Why analysis, fishbone diagrams, and fault tree analysis are three different tools built for three different kinds of failures. Used correctly and in the right order, they turn "it broke again" into a documented, permanently closed corrective action instead of another repair ticket.

Troubleshooting Fixes Today. RCFA Fixes Next Month.

Troubleshooting and root cause failure analysis solve two different problems and get confused constantly. Troubleshooting is trial and error aimed at getting equipment running again as fast as possible, and it is exactly the right response in the moment a line goes down. Root cause failure analysis is a separate, formal step that asks why the failure happened in the first place and why the systems meant to prevent it did not catch it. A plant that only troubleshoots is perpetually one failure behind, because the same root cause keeps generating new symptoms to chase. Industry data on recurring failures backs this up directly: a large share of unplanned equipment failures in manufacturing are repeat occurrences of a fault that was already repaired once on that exact asset, which means the first repair addressed the symptom and left the cause fully intact.

The confusion between the two usually comes down to timing pressure rather than a lack of understanding. In the moment a line is down, nobody is going to convene a fishbone session before restarting production, and they should not. The mistake happens afterward, when the line is back up and the investigation that should follow quietly gets skipped because the immediate crisis has passed. Building RCFA into the standard response to specific trigger events, rather than leaving it as an optional follow-up, is what keeps that second step from being the one that gets dropped every time.

Which Method Fits the Failure in Front of You

Not every failure needs the same tool. Picking the right method up front saves hours and avoids the common mistake of running a lightweight technique on a complex, multi-factor failure it was never built to handle.

5 Why Analysis
Best for: straightforward failures, single causal chain
A simple, sequential chain of "why" questions that takes fifteen to sixty minutes and needs no special tools, just disciplined questioning until the chain stops at something the team can actually fix.
Fishbone Diagram
Best for: failures with several contributing factors
A structured brainstorming map that sorts every possible cause into categories like people, machine, method, and material before the team decides which branch actually deserves further investigation.
Fault Tree Analysis
Best for: safety-critical or catastrophic failures
A top-down logic map connecting causes with AND and OR gates, used when a failure could involve multiple conditions occurring together and the stakes justify the extra rigor.

Walking a Real Failure Through the 5 Whys

The strength of the 5 Whys is that it forces the investigation past the first easy answer. Applied to the conveyor belt from the opening example, the chain of questioning typically looks something like this, with each answer becoming the next question until the chain reaches something the team can actually change.

Why 1
The belt snapped under load. Why? It was running under excessive tension.
Why 2
Why was tension excessive? The tensioning pulley had drifted out of alignment.
Why 3
Why did the pulley drift? Its mounting bolts were never checked on a scheduled interval.
Why 4
Why was there no check interval? The asset was added to the line after the PM program was built.
Why 5
Why was it missed during onboarding? New assets are not cross-checked against the PM schedule as a standard step.

Notice where the chain actually ends. It does not stop at "the pulley was misaligned," which is still a symptom. It stops at a gap in the process for onboarding new equipment into the preventive maintenance program, which is the kind of finding that prevents the next ten failures, not just this one.

Bring Your Repeat Failures

Find Out Which Asset on Your Line Is Repeating the Same Failure

If a machine has failed for the same reason more than once, the first investigation stopped short. Walk through your repeat-failure list with our team and identify where the chain actually needs to end.

The Fishbone Diagram: Six Places a Cause Can Hide

When a failure has more than one plausible cause, a sequential chain of whys is the wrong tool, because it forces the team to commit to a single path too early. A fishbone, or Ishikawa, diagram spreads the problem statement across six standard categories instead, so the team maps every plausible contributor before narrowing down to the ones actually supported by evidence.

Manpower
Training gaps, unclear responsibilities, fatigue, or a shift handover that lost information the next crew needed.
Machine
Worn components, skipped maintenance, incorrect tooling, or calibration that has drifted since the last check.
Method
A missing standard, a step performed inconsistently across shifts, or a procedure that no longer matches how the line actually runs.
Material
Out-of-spec raw material, unannounced supplier variation, a wrong batch, or damage introduced during storage.
Measurement
A miscalibrated instrument, the wrong gauge for the job, or inspection data that was recorded inconsistently.
Mother Nature
Temperature swings, humidity, ambient vibration, or dust, anything the surrounding environment does to the process.

Fault Tree Analysis: When Multiple Conditions Have to Line Up

Fault tree analysis takes a different shape entirely. Instead of asking a chain of sequential questions, it starts at an undesired top event, such as a fire, an explosion, or a catastrophic equipment failure, and works downward through logic gates until it reaches basic events that need no further breakdown. The gates are what make this method powerful for complex or safety-critical failures: an AND gate means every connected condition has to occur together for the top event to happen, while an OR gate means any single one of the connected conditions is enough on its own.

AND Gate
The top event only occurs if every input condition happens at the same time. Example: a solvent fire requires both an ignition source AND a flammable vapor concentration present simultaneously.
OR Gate
The top event occurs if any single input condition happens on its own. Example: an ignition source could come from hot work, OR an electrical fault, OR static discharge.

Because fault tree analysis can assign a probability to each basic event when reliability data exists, it also produces something the other two methods do not: a numerical estimate of how likely the top event is to occur, which is what makes it the standard choice when a failure investigation needs to support a formal risk decision rather than a single corrective action.

Choosing a Problem Statement Worth Investigating

Every method on this page depends entirely on how the initial problem statement is written, and a vague statement guarantees a vague investigation regardless of which technique follows it. "The line went down" is not a problem statement a team can trace back through causes; "the conveyor belt on Line 3 snapped at the tensioning pulley during the second shift for the third time this quarter" is. A specific, measurable statement tied to a shift, an asset, and a frequency gives every subsequent why, every fishbone branch, and every fault tree gate something concrete to test against, rather than a general impression of what went wrong.

Matching the Method to the Investigation

Lining up all three methods side by side makes it easier to choose correctly the next time a failure investigation starts, rather than defaulting to whichever technique the team happens to know best.

Method Typical Time Best Suited To Output
5 Whys 15-60 minutes Single causal chain, clear symptom One documented root cause
Fishbone Diagram 1-2 hours as a team Multiple plausible contributors Ranked list of causes by category
Fault Tree Analysis Several hours to days Safety-critical or catastrophic events Logic map with failure probability

Why Root Causes Get Found and Still Keep Failing

Finding the right cause is not the same as closing it out, and this is where most RCFA programs actually lose ground. A large majority of formal RCA investigations do successfully identify a valid root cause, yet only a fraction of those findings turn into a completed corrective action within ninety days. Once the meeting ends and the report gets filed, there is often no mechanism forcing the fix to actually happen, and the same failure mode is free to repeat while the finding sits in a folder.

No Owner Assigned
A root cause gets documented but never gets attached to a specific person and deadline, so it competes with every other open task and usually loses.
Findings Live in a Report, Not a Workflow
A PDF filed after the investigation is easy to forget. A corrective action tied directly to a work order is much harder to ignore.
No Recurrence Check
Without tracking whether the same failure mode shows up again on the same asset, nobody can tell whether the corrective action actually worked.
Different Method Every Time
When each investigation follows its own ad hoc format, there is no consistent record to compare findings against or learn from across the plant.

Building an RCFA Process That Actually Sticks

The technical part of root cause failure analysis, the whys, the fishbone, the fault tree, is rarely what makes a program fail. What determines whether an RCFA program holds up over time is the discipline wrapped around it.

Set a Clear Trigger
Define exactly which events require a formal RCFA, such as any repeat failure, any safety incident, or any downtime event above a set threshold.
Involve the People Closest to the Work
Operators and technicians who run the equipment daily surface causes that engineers and managers reviewing data alone tend to miss entirely.
Attach Every Finding to a Work Order
A corrective action with an owner, a deadline, and a linked work order gets done. A corrective action written into a report by itself often does not.
Track Recurrence, Not Just Completion
Closing the corrective action is not the finish line. Watching whether that exact failure mode returns on that asset is what proves the fix actually worked.

What Structured RCFA Programs Report

Plants that move from ad hoc troubleshooting to a structured, tracked RCFA process tend to see the improvement show up first in the assets that were previously their worst repeat offenders.

15-25%
Reduction in repeat failures
Commonly reported within the first year of a structured RCA program with consistent corrective action tracking.
Under 5%
Target recurrence rate
A widely cited benchmark for the share of RCA-investigated failures that should recur within 90 days in a mature program.
20% of causes
Drive 80% of failures
A Pareto pattern that shows up consistently once failure codes are tracked and analyzed across a plant's full asset history.
One investigation
Instead of ten repeat repairs
The intent of every method on this page: replace a cycle of repairs with a single closed, verified corrective action.

Frequently Asked Questions

Can 5 Why analysis be used for a complex, multi-factor failure?
It can be attempted, but it tends to produce a misleading single answer when a failure actually has several contributing factors happening in parallel. The sequential structure of the 5 Whys forces the team to pick one causal path early, which can bury a second or third contributing cause that never gets investigated. For failures where more than one factor plausibly contributed, a fishbone diagram is the better starting point precisely because it maps multiple branches before narrowing down. Our team can help you decide which method fits a specific failure pattern.
Who should be in the room for a fishbone diagram session?
The most effective sessions bring together the operators who run the equipment daily, the maintenance technicians who repair it, and an engineer who understands the underlying process, because each group tends to see different categories of cause. Frontline operators in particular hold detailed knowledge across the 6M categories that engineers reviewing data alone often miss. Excluding any one of these groups tends to leave entire branches of the diagram thin or missing. Book a walkthrough to see how a structured session comes together for your team.
Does every equipment failure need a formal root cause investigation?
No, running a full RCFA on every minor stoppage would consume more time than the failures themselves cost the plant. A more practical approach sets clear triggers, such as any repeat failure on the same asset, any safety-related event, or any downtime above a defined cost or duration threshold, and reserves formal investigation for those. Everything else can be logged and monitored for a pattern without a full investigation on day one. Reach out to our team to help define sensible triggers for your plant.
How is fault tree analysis different from FMEA?
Fault tree analysis is a top-down approach that starts from one specific undesired event and works backward through logic gates to find the combinations of conditions that could cause it. Failure Mode and Effects Analysis, by contrast, works bottom-up, systematically reviewing every possible failure mode of a component regardless of severity before assessing its impact. The two are complementary rather than competing, with FTA better suited to investigating one specific catastrophic scenario in depth. Talk it through with us if you are deciding which approach fits your investigation.
What is the biggest reason RCFA findings never get implemented?
The most common reason is that a valid root cause gets documented in a report but never gets converted into a tracked corrective action with an assigned owner and a deadline. Without that conversion, the finding has to compete for attention against every other open task on the floor, and it typically loses. Programs that route every RCFA finding directly into the work order system, rather than leaving it in a standalone report, see dramatically higher completion rates. Contact our team to see how findings can flow straight into your maintenance workflow.
Every Repeat Failure Is a Root Cause Nobody Closed.

Turn Your Next Investigation Into a Permanent Fix

Bring your worst repeat-failure assets to the call. We will help you pick the right method, run the investigation, and route the finding into a corrective action that actually gets tracked to completion.


Share This Story, Choose Your Platform!