Reliability-Centered Maintenance (RCM) in Power Plants

By Jackson T on October 10, 2026

power-plant-rcm-reliability-centered-maintenance

Reliability-centered maintenance (RCM) in power plants is a structured way to decide what maintenance each asset really needs. It starts from what the equipment must do, lists the ways it can fail and matches each failure to the task that deals with it best. The result is usually fewer routine overhauls, more condition checks and far better cover for hidden failures. This guide explains the method in plain language and shows how to apply it to turbines, boilers, generators and balance-of-plant rotating equipment. To see an RCM review on one of your systems, book a short walkthrough.

Power plant maintenance · RCM

Reliability-Centered Maintenance (RCM) in Power Plants: The Right Strategy for Every Asset

Not every asset needs an overhaul on a calendar. RCM matches each failure to the task that works: watch it, renew it, test it or let it run.

Quick numbers
89%
Share of items with no age-related failure pattern in the original airline study (Nowlan and Heap, 1978)
Over half
Share of coal unit forced outages caused by boiler tube leaks (NETL data via POWER)
Under 2 years
Payback reported by utilities using streamlined RCM on fossil plants (EPRI, 1998)
RCM strategy by asset group
Asset, typical failure and rcm answer
Steam turbine
Bearing wear, valve sticking
RCM answer: Condition monitoring, valve tests
Boiler
Tube leaks from wear and corrosion
RCM answer: Thickness surveys, chemistry control
Generator
Winding insulation ageing
RCM answer: Online monitoring, outage tests
Pumps and fans
Bearing and seal failure
RCM answer: Vibration and oil analysis
Protection devices
Hidden failure to operate
RCM answer: Scheduled function tests
Key takeaways
1
Start from function, not from the equipment list

RCM asks what the asset must do, then how it can fail to do it.

2
Most failures are not age-related

That is why fixed-interval overhauls often add cost without adding reliability.

3
Hidden failures need their own tasks

Standby and protective devices can fail without anyone noticing until they are needed.

4
Keep it practical

A streamlined review on critical systems delivers most of the value and actually gets implemented.

01The basics

What Is Reliability-Centered Maintenance?

RCM is a method for choosing the maintenance each asset needs, based on how it fails and what the failure costs.

In plain words
Reliability-centered maintenance

A structured review that links every maintenance task to a specific failure mode and its consequence. Tasks with no failure to prevent are removed.

  • It began in aviation. United Airlines engineers Nowlan and Heap published the method in 1978 for the US Department of Defense.
  • It reached power plants through nuclear. EPRI introduced RCM to nuclear power in 1984.
  • It has a standard. SAE JA1011 sets the criteria a process must meet to be called RCM; JA1012 is the guide.
  • It is not a software feature. It is a way of thinking that software can support.

The early results were striking. The original report notes that the DC-10 programme had 7 items subject to scheduled overhaul, against 339 on the older DC-8.

The same logic applies to a boiler feed pump or a mill. We can show it on one of your systems in a call.

02Evidence

Why Fixed-Interval Overhauls Often Miss

The airline study found that most items did not wear out on a predictable schedule.

Six failure patterns in the Nowlan and Heap study
F: early failures, then steady68%

E: steady, random14%

D: low when new, then steady7%

C: slowly rising with age5%

A: bathtub curve4%

B: wear-out at end of life2%

Share of items showing each pattern. Only patterns A, B and C, about 11% in total, are related to age.

These figures come from aircraft components, so they are not universal. Across three studies, NASA’s RCM guide reports age-related failures of 8–23%. Random failures still dominate.

If a failure is not related to age, overhauling by the calendar does not prevent it. It can even introduce early failures.

Your own failure history shows which pattern each asset follows. Our specialists can chart it.

03Method

The Seven Questions RCM Asks

Every RCM review, under SAE JA1011, works through the same seven questions in order.

1
What must it do?

The functions and performance standards, in its present operating context.

2
How can it fail to do that?

The functional failures.

3
What causes each failure?

The failure modes.

4
What happens when it fails?

The failure effects.

5
Why does it matter?

The consequences: safety, environment, output or repair cost.

6
What can be done to predict or prevent it?

The proactive tasks.

7
What if no such task exists?

The default actions: test, redesign or run to failure.

The first question matters most. A feed pump’s function is not “to run” but to deliver a set flow at a set pressure. Failing to reach that flow is a failure, even if the pump is still turning.

Working through the questions with operators and craft staff brings out knowledge no manual holds. See a worked example in a demo.

04Tasks

Matching the Task to the Failure

RCM offers a short menu of tasks. The failure mode decides which one applies.

Task typeWhen it fitsPower plant example
On-conditionThe failure gives warning before it happensVibration trend on a fan bearing
Scheduled restorationThe part wears out with age or useMill grinding element renewal by tonnes milled
Scheduled discardA life limit is knownFilter elements, elastomer seals
Failure-findingThe failure is hiddenTurbine overspeed trip test, safety valve test
Run to failureFailure is cheap and has no safety effectNon-critical indicator lamp
RedesignNo task reduces the risk enoughAdd a redundant sensor or change the material
In plain words
P-F interval

The time between the point where a developing failure can first be detected (P) and the point where the asset fails (F). Condition checks must be more frequent than this interval.

A pump bearing that gives six weeks of rising vibration before failing can be watched monthly. One that fails within days needs online monitoring or a different approach.

Task intervals come from the P-F interval, not from tradition. We set them with your data in every rollout.

05Turbines

RCM on Steam and Gas Turbines

Turbines are high-consequence assets where condition monitoring and function tests carry most of the strategy.

Failure modeConsequenceTypical RCM task
Bearing wear or damageForced outage, possible rotor damageContinuous vibration and bearing temperature monitoring
Lube oil contaminationBearing and control system damageOil sampling and analysis at set intervals
Stop or control valve stickingOverspeed risk on load rejectionRegular valve stroke tests
Overspeed protection fails to actHidden until needed; severe safety effectScheduled trip tests
Blade erosion or depositsLoss of efficiency and outputPerformance trending; inspection at outages
  • Protection first. Trip systems are hidden-failure items and need tests at defined intervals.
  • Trend performance. A slow fall in stage efficiency points to deposits or erosion.
  • Use the major outage well. Open-casing inspections confirm what the trends suggest.

Turbine overhaul intervals are increasingly set by condition and operating history. Ask our team how the evidence is assembled.

06Boilers

RCM on Boilers and Pressure Parts

Tube leaks are the largest single cause of forced outages on coal units, so the boiler is where RCM pays first.

Over half
of coal unit forced outages are boiler tube leaks
NETL data via POWER, 2021
7.1 a year
average tube leaks on wall-fired units in a 167-unit benchmark; 18.5 on cyclone units
POWER
About 3%
availability lost to tube failures on coal units above 200 MW
Inspectioneering
Failure modeHow it developsTypical RCM task
Fly ash erosionWall thins where gas speed is highThickness surveys at outages; shields
Sootblower erosionSteam jet cuts nearby tubesBlower alignment and sequence checks
Corrosion fatigue and depositsLinked to water chemistryChemistry limits and monitoring
Long-term overheatingCreep in superheater and reheater tubesMetal temperature trending; sampling
Safety valve fails to liftHidden failureScheduled valve tests

The key RCM step is to record the mechanism behind every leak, not just the repair. A tube failure reduction programme then targets each mechanism.

A tube leak history by location and mechanism is the starting point. Our engineers can help build one.

07Generators

RCM on Generators

Generator failures are rare but long and costly, so the strategy leans on online monitoring and planned tests.

Stator winding insulation
Ages with heat, vibration and electrical stress. Partial discharge monitoring tracks its condition in service.
Rotor winding
Shorted turns raise vibration and field current. Flux probe readings detect them.
Cooling system
Hydrogen purity, dew point and leakage, or stator water flow and conductivity, are checked routinely.
Bearings and seals
Vibration, temperature and seal oil condition are trended.
Excitation system
Electronic cards fail at random; spares and function checks matter more than overhaul.
Protection relays
Hidden-failure items that need scheduled secondary injection tests.

Many of these tasks are on-condition. The electrical tests that need the machine offline are grouped into planned outages.

A single view of generator health across these checks is worth having. See one in a session.

08Balance of plant

RCM on Balance-of-Plant Rotating Equipment

Pumps, fans, mills and compressors are numerous, and most give warning before they fail.

EquipmentCommon failure modesTypical RCM task
Boiler feed pumpsSeal wear, bearing damage, balance device wearVibration, seal leak-off and performance trending
ID, FD and PA fansBearing wear, blade erosion, imbalanceVibration monitoring; blade inspection at outages
Coal millsGrinding element wear, gearbox damageRenewal by throughput; gearbox oil analysis
Condensate and cooling water pumpsBearing wear, impeller erosionVibration and flow against the pump curve
Air compressorsValve and bearing wearTemperature, pressure and oil trending
Large motorsWinding insulation, bearingsCurrent analysis, thermography, insulation tests
  • Use redundancy wisely. A standby pump lowers the consequence, but only if it starts. Test it.
  • Rank by consequence. A single feed pump with no standby is critical; a spare sump pump is not.
  • Let cheap items run. Run to failure is a valid choice where the effect is small.

This is where most of the task count sits, and where trimming low-value work frees the most hours. Discuss it with our advisors.

09Hidden failures

The Failures Nobody Sees

A hidden failure is one that goes unnoticed in normal running, usually in a protective or standby device.

130

failure modes were identified in an RCM review of the coal feed system at a 600 MW coal unit. More than half were hidden, and active preventive tasks rose from 34 to 90.

Source: POWER Engineering, 1996, on the Neal 4 unit
  • Standby equipment. An auto-start that has not been tested may not work.
  • Trip and interlock devices. A stuck pressure switch is invisible until it is called on.
  • Emergency systems. DC lube oil pumps, emergency diesel sets and fire systems.

This is why RCM does not always cut the task count. It removes work that prevents nothing and adds tests where a hidden failure could turn a small event into a serious one.

A review of protective devices on one system often reveals gaps. We run one in a working session.

10Results

Does RCM Pay Back?

Published utility experience shows it does, when the analysis is kept practical and the results are implemented.

$200k–600k
annual preventive maintenance savings reported by four utilities using streamlined RCM
EPRI TR-109795, 1998
Under 1 to 2 years
payback at those utilities
EPRI TR-109795, 1998
25%
fewer maintenance labour hours across 62 systems at one nuclear plant
IAEA-TECDOC-1590

The EPRI figures cover preventive maintenance labour and parts only. Gains from avoided forced outages come on top and are harder to measure.

Results vary. The IAEA advises nuclear operators to expect a 5 to 10 year payback on a full programme, and one plant in its review found RCM cost neutral.

9.2%

weighted forced outage rate for conventional generation in North America in 2025, up from 7.6% in 2024. Coal units reached 14.1%.

Source: NERC 2026 State of Reliability

With forced outage rates rising, matching maintenance to real failure modes matters more each year. See where your units stand in a pilot.

11Approach

Classic vs Streamlined RCM

A full classic analysis of every system is rarely needed. Most fossil plants use a streamlined form.

Classic RCM
  • Every function and failure mode analysed
  • Months of team time per system
  • Best for safety-critical and new designs
  • Very thorough record
  • Risk of never finishing
Streamlined RCM
  • Focus on critical functions and known failures
  • About 152 analyst hours per system in EPRI trials
  • Suits existing fossil and combined cycle plants
  • Uses plant history and templates
  • Results in use within weeks
One 2005 survey of more than 250 companies found that over 85% of completed RCM analyses were never implemented. Keep the scope small enough to finish.

Nuclear operators often follow INPO AP-913, an equipment reliability process built on the same ideas. Our safety team can advise which approach fits.

12Steps

How to Deploy RCM in a Power Plant

Six steps take a plant from first system to routine use.

Step 1
Rank systems

By effect on safety, output and cost.

Step 2
Pick a pilot

One critical system with good history.

Step 3
Run the review

Functions, failures, consequences, tasks.

Step 4
Load the tasks

Into the CMMS with intervals and job plans.

Step 5
Measure

Failures, task hours and forced outages.

Step 6
Extend

Next system, using what was learned.

Before you start
Asset list and hierarchy in order
Failure history for the pilot system
Operators and craft staff on the team
A named owner for implementation
After the review
Every task linked to a failure mode
Hidden failures have function tests
Low-value tasks removed
Results reviewed after one year

The step most often skipped is loading the tasks. Without it the analysis stays on a shelf. Plan it with a reliability review.

13iFactory

How iFactory Supports RCM

iFactory keeps the RCM logic, the maintenance plan and the condition data in one place.

Every task traced to a failure mode, and every failure fed back into the strategy.
01
Failure mode library

Templates for common power plant equipment.

02
Criticality ranking

Systems and assets ranked by consequence.

03
Task builder

On-condition, time-based and function tests.

04
Condition monitoring

Vibration, temperature and process trends.

05
AI early warning

Models that spot drift inside the P-F interval.

06
Living programme

Each failure prompts a strategy check.

It runs on premises and links to your CMMS and historian. Share one system’s history and we will show the first results in a trial.

RCM review

See What an RCM Review Finds on One System

Share the task list and failure history for one critical system. We map tasks to failure modes and show what to keep, change or add.

Illustrative alert
Boiler feed pump BFP-2A · Strategy review

The current plan overhauls this pump every three years. Its failure history shows seal and bearing failures unrelated to age, and vibration gives six to eight weeks of warning.

Tasks linked to a failure mode58%

Window
Next planning cycle
Action
Replace the fixed overhaul with condition tasks
14In practice

A Fixed Overhaul That Was Not Preventing Failures

This is how a reliability engineer might use the system.

Reliability engineer and iFactory AI
Engineer
Feed pump 2A is due its three-year overhaul. Is that the right task?
iFactory AI
The history shows four failures in ten years: three seal failures and one bearing. None lines up with time since overhaul.
Engineer
So the overhaul is not preventing them.
iFactory AI
It does not appear to. Two of the seal failures came within four months of an overhaul, which suggests they were introduced by the work.
Engineer
What would work better?
iFactory AI
Vibration and seal leak-off trends gave six to eight weeks of warning each time. I suggest monthly condition checks and overhaul on condition, and keeping the yearly auto-start test on the standby pump.
Turnkey hardware and software

iFactory ships as a pre-configured NVIDIA AI server, racked and ready with the maintenance and reliability models loaded. Rack it, plug in power and Ethernet, and the AI is live. Scope covers data connections across units, control room, stores and planning office, DCS, historian, CMMS and ERP integration, cabling and network setup, team training and 24×7 remote monitoring.

Weeks 1–4
Ship, network, data

Server installed, DCS, historian and CMMS links live, history loaded.

Weeks 5–8
Train models, pilot

Models tuned on your own plant data, then piloted on one unit with your team reviewing every output.

Weeks 9–12
Go live, train teams

Rollout to the agreed units, team training done, 24×7 remote monitoring in place.

Software, server and integration come as one package. For pricing, contact our sales team.

FAQQuestions

Frequently Asked Questions

What is RCM in a power plant?

A structured method for deciding the maintenance each asset needs. It links every task to a failure mode and its consequence, and removes tasks that prevent nothing.

What are the seven questions of RCM?

Functions, functional failures, failure modes, failure effects, consequences, proactive tasks and default actions, as set out in SAE JA1011.

Does RCM reduce maintenance work?

Often, but not always. It usually removes low-value overhauls and adds condition checks and tests for hidden failures. The aim is the right work, not less work.

Which power plant systems should be reviewed first?

Start with systems that cause the most forced outages. On coal units that is usually the boiler, since tube leaks cause over half of forced outages.

What is the difference between classic and streamlined RCM?

Classic RCM analyses every function and failure mode. Streamlined RCM focuses on critical functions and known failures, and is faster to complete and implement.

How long does an RCM pilot take?

A review of one critical system, with tasks loaded and monitoring live, typically fits within a 6–12 week rollout. Plan it with our specialists.

Next step

Give Every Asset the Maintenance It Needs

iFactory links failure modes, tasks and condition data, so your maintenance plan follows how equipment really fails.

Illustrative dashboard view
Maintenance tasks by type after an RCM review
Condition-based42%

Time-based26%

Failure-finding14%

Run to failure12%

Redesign6%

Illustrative. The mix differs by plant; the point is that every task traces to a failure mode.


Share This Story, Choose Your Platform!