Equipment Reliability Improvement: Bad Actor Elimination Tips

By James Smith on August 8, 2026

equipment-reliability-improvement-program-bad-actor

Reliability engineers have a name for what happens on almost every plant floor: 10 to 20 percent of installed equipment generates 80 to 90 percent of total downtime, repair cost, and lost production — and most maintenance teams keep fixing that same short list of machines over and over without ever asking why they keep breaking. The Society for Maintenance & Reliability Professionals estimates that eliminating chronic failures like these can cut overall maintenance costs by up to 60 percent, because they represent the majority of hidden factory waste. iFactory's reliability improvement platform turns scattered failure history into a live, ranked list of exactly which assets deserve engineering attention first.

Equipment Failures → Bad Actor Elimination

10-20% of Your Equipment Is Quietly Draining 80% of Your Maintenance Budget

A bad actor is not the machine that broke down dramatically last week. It is the machine that keeps breaking down quietly, month after month, until everyone treats its failures as normal. iFactory finds these assets automatically and ranks them by real financial impact.

Typical Plant: Downtime Hours by Asset Rank
Vital Few Trivial Many
First 3-4 assets (left) account for the majority of total downtime hours — the rest is long-tail noise
The Elimination Program

Five Stages That Turn a Chronic Failure Into a Solved Problem

Bad actor elimination is not a single root cause analysis meeting. It is a narrowing funnel — starting with every asset in the plant and ending with a permanent fix for the handful that matter most. Skipping a stage is how the same failure keeps coming back with a new work order number. Plants that treat elimination as a one-time project rather than a running funnel tend to solve one bad actor only to discover, six months later, that a different asset has quietly taken its place at the top of the list.

1
Detect
Pull 12-24 months of work order history across every asset. Calculate MTBF, MTTR, and total cost of unreliability for each one. This is the stage most plants skip or shortcut, relying instead on which machine generated the loudest complaint this month rather than which one actually cost the most over the full year.
2
Rank
Run Pareto analysis on downtime hours and repair cost. Cross-reference against asset criticality — a bad actor on a bottleneck machine outranks one on a spared, non-critical line. Two assets with identical failure counts can carry wildly different consequences depending on where they sit in the process, which is why frequency alone is never a sufficient ranking signal on its own.
3
Diagnose
Run root cause analysis on the top-ranked assets using 5-Whys, fishbone, or fault tree analysis. Stop at the true source, not the first symptom. A single failure mode usually needs only one RCA, but an asset with multiple distinct failure modes over the review period needs a separate investigation for each one — combining them produces a muddled, unactionable conclusion.
4
Eliminate
Deploy the corrective action — a design change, a procedure fix, a material upgrade, or a maintenance strategy shift — matched to the actual root cause found in Stage 3. Cross-functional input matters here: an operator, a maintenance technician, and a reliability engineer often see different parts of the same failure pattern, and the strongest corrective actions usually combine all three perspectives.
5
Verify
Track MTBF and downtime cost on the treated asset for 90+ days. If the failure returns, the RCA missed the real root cause — reopen the investigation.

Chronic vs Sporadic: The Distinction That Decides Your Strategy

Not every failure is a bad actor, and treating them the same way wastes engineering time. A sporadic failure is a single dramatic event with an obvious cause — a forklift striking a panel, a power surge. A chronic failure is quiet, repetitive, and gets normalized by the people who deal with it every week. The danger is not that chronic failures go unnoticed — operators and technicians usually know exactly which machine is the problem child — it is that nobody ever formally escalates that knowledge into a ranked, resourced investigation.

Sporadic Failure
PatternSingle, dramatic, isolated event
Root CauseUsually one obvious, external cause
Team ResponseFix and move on — appropriate response
Correct ProgramStandard corrective work order
Chronic Failure — The Bad Actor
PatternFrequent, repetitive, often normalized
Root CauseMultiple contributing factors — lubrication, procedure, design
Team ResponseRepeated patch repairs that never resolve the pattern
Correct ProgramFormal RCA with tracked corrective action to closure
Diagnostic Toolkit

Three Root Cause Methods and When Each One Actually Fits

Not every chronic failure needs the same investigative depth. Choosing a method that is heavier than the problem wastes engineering time; choosing one that is lighter than the problem produces a shallow answer that will not survive contact with the next failure. Reliability teams generally reach for one of three tools depending on how the failure presents, and mixing them incorrectly is a common source of RCA reports that read as thorough but never actually change anything on the floor.

5-Whys
Best fit for failures with a fairly linear cause-and-effect chain. The team asks "why" repeatedly against the previous answer until the trail stops at something actionable — a procedure, a spec, a training gap — rather than another symptom. Fast to run, but weak on failures with multiple contributing factors happening simultaneously.
Fishbone (Ishikawa) Diagram
Best fit when a failure likely has several contributing categories at once — method, machine, material, manpower, environment. The visual structure forces the team to consider categories they might otherwise skip, which is particularly useful for failures that have resisted a single-cause explanation in previous attempts.
Fault Tree Analysis (FTA)
Best fit for safety-critical or high-consequence failures where understanding every combination of contributing conditions matters, not just the most likely path. More time-intensive than the other two methods, which is why it is typically reserved for the highest-criticality assets on the ranked list rather than applied universally.
Before / After

What MTBF Improvement Actually Looks Like on a Treated Asset

Mean Time Between Failures is the clearest proof that a root cause elimination effort worked. Reliability-centered maintenance programs that follow the detect-rank-diagnose-eliminate-verify sequence consistently show meaningful MTBF gains within the first two quarters.

Rotary/Continuous Equipment
Before: 1,000 hrs avg
After: 1,400 hrs avg (+40%)
Rotating Assets (Pumps, Motors)
Before baseline
After: +33% MTBF
RCA-Driven Interventions (General)
Before baseline
After: +29-38% MTBF
Ranges reflect documented outcomes across reliability-centered maintenance programs following structured RCA; actual results vary by asset class, baseline maturity, and root cause category.
Program Ownership

Who Actually Runs a Bad Actor Program Day to Day

Bad actor elimination fails when it is treated as one person's side project. It works best as a shared responsibility with clear handoffs between three roles that rarely report to the same manager but each hold a piece of the picture no one else has.

Operators
First line of observation. They notice the early warning signs — an unusual sound, a temperature shift, a subtle change in cycle time — long before a failure shows up in the CMMS. Their informal knowledge of "which machine is always trouble" is often the fastest way to shortlist candidates before running the formal Pareto numbers, provided there is a clear channel for that observation to actually reach the reliability team instead of staying as shift-change conversation.
Maintenance Technicians
Hold the repair history that a spreadsheet cannot capture — which parts actually failed, what the failure looked like on teardown, and whether the same symptom has shown up before under a different work order. Their input is essential during the diagnose stage of the funnel, not just the eliminate stage.
Reliability Engineers
Own the ranking methodology, run or facilitate the formal RCA, and track verification data after a fix is deployed. In plants without a dedicated reliability role, this responsibility typically falls to a maintenance manager or a designated bad actor champion who coordinates across the other two roles — the title matters far less than making sure someone is explicitly accountable for closing the loop from ranking through verification.

Stop Re-Diagnosing the Same Failure Every Quarter

If a work order description keeps repeating on the same asset tag, the root cause was never actually found. iFactory surfaces the pattern automatically, before the fifth repair becomes the sixth.

Identification Criteria

How to Know an Asset Actually Qualifies as a Bad Actor

Not every troublesome machine is a bad actor in the formal sense, and misclassifying assets wastes engineering hours on the wrong targets. Reliability programs typically score candidates against four weighted criteria before committing RCA resources. A weighted scoring approach — rather than ranking on any single criterion alone — is what prevents a loud but low-cost nuisance machine from crowding out a quieter, far more expensive failure elsewhere in the plant.

Frequency
Failure Rate vs Asset Class Norm
For spared, general-purpose machinery, a common threshold is one or more failures per year. For critical, unspared machinery, even a single failure between planned turnarounds can qualify — the acceptable bar is much lower. The distinction matters because applying a single failure-frequency threshold across every asset class treats a backup conveyor motor the same as a single-train compressor, which is rarely the right comparison.
Cost Impact
Cost of Unreliability (CoUR)
Combined repair cost, lost production value, and expedited parts or labor premiums over a rolling 12-month window. This is usually the single largest ranking factor in a weighted Bad Actor Index, since it converts every failure — regardless of how technically interesting or mundane it looks on paper — into the same comparable unit that budget conversations actually run on.
Criticality
Position in the Process
A chronic failure on a bottleneck or single-point-of-failure asset carries systemic risk that the same failure rate on a redundant, non-critical asset does not. Criticality ranking filters which failures deserve engineering time first, and it is the criterion most often skipped by teams that default to a simple frequency-only Pareto chart.
Pattern
Repeated or Related Failure Mode
The same or a closely related failure mode showing up across multiple repair records — not just a rising count of unrelated issues — is the signature of a true bad actor rather than an asset that is simply aging normally. Distinguishing pattern from coincidence usually requires reviewing the actual repair notes and failure descriptions together, not just tallying work order counts.
Where Programs Break Down

The Recurring Ways Bad Actor Programs Quietly Fail

A reliability program rarely collapses all at once. It erodes through small, individually reasonable shortcuts that compound over several quarters until the ranked list stops reflecting reality. The four patterns below account for most of the programs that start strong and quietly stall out.

01
The RCA stops at the first symptom.
"Bearing failure" is a symptom, not a root cause. If the RCA does not trace back to why the bearing failed — contamination, misalignment, wrong lubricant — the same asset returns to the list within a quarter under a different work order number. Teams under schedule pressure are especially prone to this shortcut, because a symptom-level fix looks identical to a root-cause fix in the short term and only reveals the difference months later.
02
Work orders get closed without a failure code.
A Pareto ranking is only as good as the data behind it. When technicians close corrective work orders without assigning a failure code, the ranking silently degrades and the wrong assets rise to the top of the priority list. This is one of the quietest ways a program fails, because the ranking still produces output — it simply stops being trustworthy without anyone noticing until the numbers are audited against reality.
03
Every failure gets treated with equal urgency.
Without a ranked list, maintenance teams spread limited engineering hours evenly across every complaint. The vital few bad actors responsible for most of the downtime never receive the concentrated attention that would actually fix them.
04
Nobody tracks the asset after the fix.
A corrective action is a hypothesis until it is verified. Programs that skip post-fix MTBF tracking cannot tell the difference between a genuine root cause elimination and a repair that happened to hold for a few extra weeks. Verification is the least glamorous stage of the funnel, which is exactly why it is the one most often skipped once the team moves on to the next asset on the list.
Financial Case

Why Bad Actor Elimination Consistently Pays for Itself

The financial case for bad actor elimination rarely needs a sophisticated model to make. Reactive, unplanned repair work costs roughly four times as much as the same repair done on a planned schedule, and a small number of assets are responsible for the overwhelming majority of that reactive spend. The numbers below reflect what changes once a plant starts treating that concentration as a targeting opportunity rather than background noise.

What Chronic Failures Actually Cost
Cost of unplanned/emergency work vs planned work~4x higher
Reduction in maintenance cost from eliminating chronic failuresUp to 60%
Downtime concentration in top 3 failure causesOften >70%
Typical unplanned downtime reduction from targeted programs30-50%
What a Ranked Program Changes
Engineering hours concentrate on the 3-5 assets responsible for most of the downtime, instead of spreading thin across every complaint on the floor.
MTBF improvement on treated assets becomes measurable and verifiable within 90 days, giving reliability teams a defensible number for budget conversations.
New bad actors are caught by the ranking system as soon as their failure pattern crosses threshold — not months later when someone notices the repair invoices piling up.
"

The plants that never make progress on bad actors are almost never short on effort — they are short on ranking discipline. Everyone knows the pump that fails every six weeks. What they miss is that a less dramatic but far more expensive failure is quietly costing three times as much down the hall, because nobody ever ran the Pareto numbers to compare them side by side. The moment a plant starts ranking failures by actual dollar impact instead of by how loud the complaint was, the priority list changes completely — and that is usually the first real progress a reliability program makes.

Marcus Whitfield
Reliability Engineering Manager — 21 Years in Discrete & Process Manufacturing
Common Questions

Bad Actor Elimination — Frequently Asked Questions

How is a bad actor different from just an old or unreliable machine?
Age alone does not make an asset a bad actor. The defining characteristic is that its failure frequency and associated cost exceed what is normal and acceptable for its asset class — an old machine performing in line with its peer group is not a bad actor, while a relatively new machine failing repeatedly due to a design flaw or improper installation absolutely is. iFactory's reliability dashboard benchmarks each asset against its own class, not against an arbitrary age cutoff, so genuine outliers surface regardless of how old the equipment happens to be.
How many assets should a plant expect to have on its bad actor list at any given time?
Following the Pareto principle, most plants find that roughly 10-20% of their critical assets qualify as active bad actors at any point, though the specific number depends heavily on plant age, process complexity, and how mature the existing maintenance program already is. A newly launched program often surfaces more candidates simply because nobody has previously ranked failures by financial impact; that list should shrink measurably as root causes get eliminated over successive quarters.
What happens if the same bad actor keeps returning after a fix has already been implemented?
A returning bad actor almost always means the root cause analysis stopped at a symptom rather than the true underlying cause — for example, replacing a failed bearing repeatedly without ever investigating why the bearing keeps failing, such as a contaminated lubrication source or a misalignment issue upstream. The correct response is to reopen the investigation with a broader cross-functional team and preserved failure evidence from the asset, rather than simply scheduling another patch repair. Preserving physical evidence from the failed component — rather than discarding it immediately during the repair — often makes the difference between a second RCA that finds the real cause and one that repeats the same guesswork as the first attempt. Book a demo to see how iFactory flags repeat failure patterns automatically.
Can a small plant with a limited maintenance team realistically run a bad actor elimination program?
Yes — the methodology scales down cleanly because the core discipline is prioritization, not headcount. A two-person maintenance team gains more from ranking its worst three assets and fixing them properly than from spreading the same effort evenly across twenty complaints. Starting with a simple Pareto ranking of existing work order history, even without dedicated reliability engineering staff, typically reveals the highest-impact starting point within the first review, and the same ranked-list discipline scales up naturally as the team and the plant both grow.
How soon should a plant expect to see measurable results after starting a bad actor elimination program?
Most programs see measurable MTBF improvement on treated assets within six months, with broader forced outage rate reductions in the 20-40% range becoming visible within 18 months as more of the ranked list gets addressed. Results compound over time because the ranking and verification steps are continuous rather than one-time — new bad actors will always emerge as older ones get resolved, and the program's job is to keep catching them early rather than reaching a finish line.

Find Your Plant's Vital Few Before They Cost You Another Quarter

iFactory ranks every asset by real financial impact, tracks failure patterns automatically, and verifies whether a fix actually worked — so engineering time goes to the machines that matter most, and the same failure never quietly reclaims its spot on next quarter's repair invoice.


Share This Story, Choose Your Platform!