Reliability-centered maintenance (RCM) in power plants is a structured way to decide what maintenance each asset really needs. It starts from what the equipment must do, lists the ways it can fail and matches each failure to the task that deals with it best. The result is usually fewer routine overhauls, more condition checks and far better cover for hidden failures. This guide explains the method in plain language and shows how to apply it to turbines, boilers, generators and balance-of-plant rotating equipment. To see an RCM review on one of your systems, book a short walkthrough.
Reliability-Centered Maintenance (RCM) in Power Plants: The Right Strategy for Every Asset
Not every asset needs an overhaul on a calendar. RCM matches each failure to the task that works: watch it, renew it, test it or let it run.
RCM asks what the asset must do, then how it can fail to do it.
That is why fixed-interval overhauls often add cost without adding reliability.
Standby and protective devices can fail without anyone noticing until they are needed.
A streamlined review on critical systems delivers most of the value and actually gets implemented.
What Is Reliability-Centered Maintenance?
RCM is a method for choosing the maintenance each asset needs, based on how it fails and what the failure costs.
A structured review that links every maintenance task to a specific failure mode and its consequence. Tasks with no failure to prevent are removed.
- It began in aviation. United Airlines engineers Nowlan and Heap published the method in 1978 for the US Department of Defense.
- It reached power plants through nuclear. EPRI introduced RCM to nuclear power in 1984.
- It has a standard. SAE JA1011 sets the criteria a process must meet to be called RCM; JA1012 is the guide.
- It is not a software feature. It is a way of thinking that software can support.
The early results were striking. The original report notes that the DC-10 programme had 7 items subject to scheduled overhaul, against 339 on the older DC-8.
The same logic applies to a boiler feed pump or a mill. We can show it on one of your systems in a call.
Why Fixed-Interval Overhauls Often Miss
The airline study found that most items did not wear out on a predictable schedule.
Share of items showing each pattern. Only patterns A, B and C, about 11% in total, are related to age.
These figures come from aircraft components, so they are not universal. Across three studies, NASA’s RCM guide reports age-related failures of 8–23%. Random failures still dominate.
Your own failure history shows which pattern each asset follows. Our specialists can chart it.
The Seven Questions RCM Asks
Every RCM review, under SAE JA1011, works through the same seven questions in order.
The functions and performance standards, in its present operating context.
The functional failures.
The failure modes.
The failure effects.
The consequences: safety, environment, output or repair cost.
The proactive tasks.
The default actions: test, redesign or run to failure.
The first question matters most. A feed pump’s function is not “to run” but to deliver a set flow at a set pressure. Failing to reach that flow is a failure, even if the pump is still turning.
Working through the questions with operators and craft staff brings out knowledge no manual holds. See a worked example in a demo.
Matching the Task to the Failure
RCM offers a short menu of tasks. The failure mode decides which one applies.
| Task type | When it fits | Power plant example |
|---|---|---|
| On-condition | The failure gives warning before it happens | Vibration trend on a fan bearing |
| Scheduled restoration | The part wears out with age or use | Mill grinding element renewal by tonnes milled |
| Scheduled discard | A life limit is known | Filter elements, elastomer seals |
| Failure-finding | The failure is hidden | Turbine overspeed trip test, safety valve test |
| Run to failure | Failure is cheap and has no safety effect | Non-critical indicator lamp |
| Redesign | No task reduces the risk enough | Add a redundant sensor or change the material |
The time between the point where a developing failure can first be detected (P) and the point where the asset fails (F). Condition checks must be more frequent than this interval.
A pump bearing that gives six weeks of rising vibration before failing can be watched monthly. One that fails within days needs online monitoring or a different approach.
Task intervals come from the P-F interval, not from tradition. We set them with your data in every rollout.
RCM on Steam and Gas Turbines
Turbines are high-consequence assets where condition monitoring and function tests carry most of the strategy.
| Failure mode | Consequence | Typical RCM task |
|---|---|---|
| Bearing wear or damage | Forced outage, possible rotor damage | Continuous vibration and bearing temperature monitoring |
| Lube oil contamination | Bearing and control system damage | Oil sampling and analysis at set intervals |
| Stop or control valve sticking | Overspeed risk on load rejection | Regular valve stroke tests |
| Overspeed protection fails to act | Hidden until needed; severe safety effect | Scheduled trip tests |
| Blade erosion or deposits | Loss of efficiency and output | Performance trending; inspection at outages |
- Protection first. Trip systems are hidden-failure items and need tests at defined intervals.
- Trend performance. A slow fall in stage efficiency points to deposits or erosion.
- Use the major outage well. Open-casing inspections confirm what the trends suggest.
Turbine overhaul intervals are increasingly set by condition and operating history. Ask our team how the evidence is assembled.
RCM on Boilers and Pressure Parts
Tube leaks are the largest single cause of forced outages on coal units, so the boiler is where RCM pays first.
| Failure mode | How it develops | Typical RCM task |
|---|---|---|
| Fly ash erosion | Wall thins where gas speed is high | Thickness surveys at outages; shields |
| Sootblower erosion | Steam jet cuts nearby tubes | Blower alignment and sequence checks |
| Corrosion fatigue and deposits | Linked to water chemistry | Chemistry limits and monitoring |
| Long-term overheating | Creep in superheater and reheater tubes | Metal temperature trending; sampling |
| Safety valve fails to lift | Hidden failure | Scheduled valve tests |
The key RCM step is to record the mechanism behind every leak, not just the repair. A tube failure reduction programme then targets each mechanism.
A tube leak history by location and mechanism is the starting point. Our engineers can help build one.
RCM on Generators
Generator failures are rare but long and costly, so the strategy leans on online monitoring and planned tests.
Many of these tasks are on-condition. The electrical tests that need the machine offline are grouped into planned outages.
A single view of generator health across these checks is worth having. See one in a session.
RCM on Balance-of-Plant Rotating Equipment
Pumps, fans, mills and compressors are numerous, and most give warning before they fail.
| Equipment | Common failure modes | Typical RCM task |
|---|---|---|
| Boiler feed pumps | Seal wear, bearing damage, balance device wear | Vibration, seal leak-off and performance trending |
| ID, FD and PA fans | Bearing wear, blade erosion, imbalance | Vibration monitoring; blade inspection at outages |
| Coal mills | Grinding element wear, gearbox damage | Renewal by throughput; gearbox oil analysis |
| Condensate and cooling water pumps | Bearing wear, impeller erosion | Vibration and flow against the pump curve |
| Air compressors | Valve and bearing wear | Temperature, pressure and oil trending |
| Large motors | Winding insulation, bearings | Current analysis, thermography, insulation tests |
- Use redundancy wisely. A standby pump lowers the consequence, but only if it starts. Test it.
- Rank by consequence. A single feed pump with no standby is critical; a spare sump pump is not.
- Let cheap items run. Run to failure is a valid choice where the effect is small.
This is where most of the task count sits, and where trimming low-value work frees the most hours. Discuss it with our advisors.
The Failures Nobody Sees
A hidden failure is one that goes unnoticed in normal running, usually in a protective or standby device.
failure modes were identified in an RCM review of the coal feed system at a 600 MW coal unit. More than half were hidden, and active preventive tasks rose from 34 to 90.
- Standby equipment. An auto-start that has not been tested may not work.
- Trip and interlock devices. A stuck pressure switch is invisible until it is called on.
- Emergency systems. DC lube oil pumps, emergency diesel sets and fire systems.
This is why RCM does not always cut the task count. It removes work that prevents nothing and adds tests where a hidden failure could turn a small event into a serious one.
A review of protective devices on one system often reveals gaps. We run one in a working session.
Does RCM Pay Back?
Published utility experience shows it does, when the analysis is kept practical and the results are implemented.
The EPRI figures cover preventive maintenance labour and parts only. Gains from avoided forced outages come on top and are harder to measure.
Results vary. The IAEA advises nuclear operators to expect a 5 to 10 year payback on a full programme, and one plant in its review found RCM cost neutral.
weighted forced outage rate for conventional generation in North America in 2025, up from 7.6% in 2024. Coal units reached 14.1%.
With forced outage rates rising, matching maintenance to real failure modes matters more each year. See where your units stand in a pilot.
Classic vs Streamlined RCM
A full classic analysis of every system is rarely needed. Most fossil plants use a streamlined form.
- Every function and failure mode analysed
- Months of team time per system
- Best for safety-critical and new designs
- Very thorough record
- Risk of never finishing
- Focus on critical functions and known failures
- About 152 analyst hours per system in EPRI trials
- Suits existing fossil and combined cycle plants
- Uses plant history and templates
- Results in use within weeks
Nuclear operators often follow INPO AP-913, an equipment reliability process built on the same ideas. Our safety team can advise which approach fits.
How to Deploy RCM in a Power Plant
Six steps take a plant from first system to routine use.
By effect on safety, output and cost.
One critical system with good history.
Functions, failures, consequences, tasks.
Into the CMMS with intervals and job plans.
Failures, task hours and forced outages.
Next system, using what was learned.
The step most often skipped is loading the tasks. Without it the analysis stays on a shelf. Plan it with a reliability review.
How iFactory Supports RCM
iFactory keeps the RCM logic, the maintenance plan and the condition data in one place.
Templates for common power plant equipment.
Systems and assets ranked by consequence.
On-condition, time-based and function tests.
Vibration, temperature and process trends.
Models that spot drift inside the P-F interval.
Each failure prompts a strategy check.
It runs on premises and links to your CMMS and historian. Share one system’s history and we will show the first results in a trial.
See What an RCM Review Finds on One System
Share the task list and failure history for one critical system. We map tasks to failure modes and show what to keep, change or add.
The current plan overhauls this pump every three years. Its failure history shows seal and bearing failures unrelated to age, and vibration gives six to eight weeks of warning.
A Fixed Overhaul That Was Not Preventing Failures
This is how a reliability engineer might use the system.
iFactory ships as a pre-configured NVIDIA AI server, racked and ready with the maintenance and reliability models loaded. Rack it, plug in power and Ethernet, and the AI is live. Scope covers data connections across units, control room, stores and planning office, DCS, historian, CMMS and ERP integration, cabling and network setup, team training and 24×7 remote monitoring.
Server installed, DCS, historian and CMMS links live, history loaded.
Models tuned on your own plant data, then piloted on one unit with your team reviewing every output.
Rollout to the agreed units, team training done, 24×7 remote monitoring in place.
Software, server and integration come as one package. For pricing, contact our sales team.
Frequently Asked Questions
A structured method for deciding the maintenance each asset needs. It links every task to a failure mode and its consequence, and removes tasks that prevent nothing.
Functions, functional failures, failure modes, failure effects, consequences, proactive tasks and default actions, as set out in SAE JA1011.
Often, but not always. It usually removes low-value overhauls and adds condition checks and tests for hidden failures. The aim is the right work, not less work.
Start with systems that cause the most forced outages. On coal units that is usually the boiler, since tube leaks cause over half of forced outages.
Classic RCM analyses every function and failure mode. Streamlined RCM focuses on critical functions and known failures, and is faster to complete and implement.
A review of one critical system, with tasks loaded and monitoring live, typically fits within a 6–12 week rollout. Plan it with our specialists.
Give Every Asset the Maintenance It Needs
iFactory links failure modes, tasks and condition data, so your maintenance plan follows how equipment really fails.
Illustrative. The mix differs by plant; the point is that every task traces to a failure mode.







