Power Plant Deviation Investigation and RCA Best Practices

By David Cook on October 10, 2026

power-plant-deviation-investigation-rca

Power plant deviation investigation and root cause analysis (RCA) turn a trip, a chemistry excursion or a limit breach into a fix that stops it happening again. Too often the event is logged, a report is filed and the unit goes back on line, only for the same sequence to repeat months later. A structured model changes that: capture the evidence, find the real cause, act on it and check that the action worked. This guide sets out each step with the standards behind it. To see a closed-loop investigation on one of your events, book a short walkthrough.

Power plant quality · Deviation and RCA

Power Plant Deviation Investigation and RCA Best Practices: Close the Loop on Repeat Trips

From first evidence to verified fix. A clear model for investigating trips, chemistry excursions and limit breaches so the same event does not return.

Quick numbers
5 phases
In the US Department of Energy root cause analysis guidance
0.5–1 ms
Typical resolution of a sequence-of-events recorder, against seconds for a historian
3.6%
Share of bulk power event root causes attributed to individual human performance (NERC data)
Five phases of an investigation
Phase and what happensWhen
Collect
First hours
Secure event data and statements at once
Assess
Days
Find direct, contributing and root causes
Correct
Days to weeks
Actions that remove each cause
Inform
On completion
Share findings with other units
Follow up
Months later
Check the actions worked
Key takeaways
1
Restarting is not resolving

A unit back on line with the cause unknown is waiting for the next trip.

2
Evidence fades fast

Secure sequence-of-events data, trends and statements in the first hour.

3
Look past the operator

Industry data shows design and organisational causes far outnumber individual error.

4
An action is not closed until it is proven

The effectiveness check is what breaks the repeat cycle.

01The basics

What Counts as a Deviation?

A deviation is any event where the plant, a process or a product moved outside its defined limits.

Trip
Unit or equipment

A protective shutdown, planned or not.

Chemistry
Water and steam

A parameter outside its limit for longer than allowed.

Limit breach
Operating envelope

Temperature, pressure, vibration or emission above a set value.

Quality
Fuel and consumables

A delivery or sample that fails specification.

Near miss
No consequence yet

An event that could have caused a trip or injury.

Repeat
Seen before

Any event that matches an earlier one.

In plain words
Root cause

In the words of US Department of Energy guidance: “the cause that, if corrected, would prevent recurrence of this and similar occurrences.”

That last phrase matters. A good investigation prevents similar events on other equipment too, not only the one that failed.

One clear definition of what gets investigated is the starting point. We can review yours on a call.

02The problem

Why the Same Events Come Back

After a trip, pressure to restart is high. Finding the cause can lose out to getting back on line.

Over half
of coal unit forced outages are boiler tube leaks
NETL data via POWER, 2021
About 3%
availability lost to tube failures on coal units above 200 MW
Inspectioneering
8 starts
of cyclic life used by one gas turbine trip from full load
TWI
  • Repair without mechanism. The tube is replaced; why it failed is never established.
  • Report without action. The form is filed, and nothing in the plant changes.
  • Action without check. A fix is made and closed, and nobody confirms it worked.
  • No link between events. Each trip is treated as new, so patterns are missed.

One industry review notes that tube failures “are usually repetitive in nature and result in multiple forced outages”, and that return to service has historically come before finding the mechanism.

We found no published statistic on how often trips repeat. Your own event history is the place to measure it.

A count of repeat events in your last two years is a revealing first exercise. See how in a demo.

03The model

A Five-Phase Investigation Model

US Department of Energy guidance sets out five phases that fit any power plant event.

Step 1
Collect

Data, trends, alarms, statements, physical evidence.

Step 2
Assess

Analyse to find the causes.

Step 3
Correct

Define and carry out actions.

Step 4
Inform

Report and share the lessons.

Step 5
Follow up

Confirm the actions were effective.

PhaseKey questionOutput
CollectWhat exactly happened, in what order?Timeline and evidence file
AssessWhy did it happen?Direct, contributing and root causes
CorrectWhat will stop it happening again?Actions with owners and dates
InformWho else needs to know?Report and lessons for other units
Follow upDid it work?Effectiveness review

The source is DOE-NE-STD-1004-92, a root cause analysis guidance document written for nuclear facilities and widely borrowed since.

Most plants do the first three phases. The last two are where loops stay open. Our specialists can show a full cycle.

04Evidence

Secure the Evidence in the First Hour

The order of events decides the cause, and only some systems record it precisely enough.

0.5–1 ms
typical resolution of a sequence-of-events recorder
Industrial Monitor Direct
5–20 ms
typical controller scan time
Industrial Monitor Direct
1–10 s
typical historian sampling interval
Industrial Monitor Direct
In plain words
First-out

Logic that identifies which of several near-simultaneous trip signals came first. It points to the initiating event.

  • Sequence-of-events printout. Save it before it is overwritten.
  • Trends. Key parameters for at least an hour before the event.
  • Alarm and event list. Including alarms that were standing beforehand.
  • Statements. From each operator, written separately and soon.
  • Physical evidence. Failed parts, photographs, samples.
  • Plant state. Line-ups, permits, work in progress, bypasses in force.

A historian sampling every few seconds cannot tell which of two signals a few milliseconds apart came first. The sequence-of-events record can.

Automatic capture of this data at every trip removes the scramble. We set it up in every rollout.

05Methods

Choosing an Analysis Method

Different methods suit different events. Simple events need simple tools.

MethodHow it worksBest for
5 WhysAsk why repeatedly until the underlying cause appearsSimple events with one causal chain
Fishbone diagramGroup possible causes under headings such as machine, method and materialBrainstorming when the cause is unclear
Events and causal factors chartTimeline with conditions and causes attachedTrips with a sequence of steps
Change analysisCompare this event with a time it went rightEvents after a modification or new procedure
Barrier analysisIdentify the defences that failed or were missingSafety and protection events
Fault tree analysisWork down from the top event through logic gatesComplex or high-consequence events
5 Whys
Developed by Sakichi Toyoda and used at Toyota. Critics note that investigators can stop at symptoms.
Fishbone
Popularised by Kaoru Ishikawa in the 1960s.
Fault tree
First used in 1962 at Bell Labs; covered by standard IEC 61025.

For most trips, a timeline chart plus 5 Whys on each causal factor is enough. Ask our team for worked examples.

06Causes

Look Beyond “Operator Error”

Stopping at human error is the most common reason corrective actions fail.

3.6%

of root causes in 1,007 bulk power system events from 2010 to 2023 were individual human performance. Organisational causes were 41% and design or engineering 26.4%.

Source: NERC event analysis data, presented by WECC
Cause categoryExample in a power plant
Equipment or materialFailed transmitter, worn seal
ProcedureStep missing or unclear in the start-up procedure
PersonnelWrong valve operated
DesignSingle sensor with no redundancy on a trip
TrainingOperator never practised the scenario
ManagementKnown defect deferred without risk review
ExternalGrid disturbance, fuel quality, weather

The seven categories are from the Department of Energy guidance. If the answer is “personnel”, ask one more question: what in the system made that error likely?

Coding every event to one of these lets you see patterns across years. Our engineers can set up the coding.

07Chemistry

Chemistry Deviations and Action Levels

Chemistry excursions rarely trip a unit, but they cause the tube and turbine damage that leads to later outages.

In plain words
Action level

A band above the normal limit for a chemistry parameter. Each band sets how long the plant may keep running before it must act.

LevelMeaningExample response
NormalWithin targetRoutine monitoring
Action level 1Outside targetFind and correct the cause; escalates if not restored in time
Action level 2Further outsideShorter time allowed; load reduction if not restored
Action level 3SevereImmediate shutdown

In one US nuclear plant procedure for secondary water, Level 1 escalates after a week, Level 2 requires a power reduction after eight hours and Level 3 requires immediate shutdown.

For fossil and combined cycle plants, IAPWS guidance suggests a first action level at twice the target value and a second at four times. Time limits come from each plant’s own chemistry guidelines.

  • Record time outside limits. Hours at each level, not just the peak value.
  • Investigate like a trip. Condenser leak, dosing fault, sampling error or resin exhaustion.
  • Link to later failures. A tube leak investigation should always check chemistry history.

Tracking cumulative hours at each level turns chemistry into a leading indicator. See it in a session.

08Triage

Which Events Need a Full RCA?

Not every deviation deserves a team and a week. Grade the effort to the consequence.

GradeTypical eventsInvestigation
A: significantUnit trip with damage, safety event, repeat trip, action level 3Full root cause analysis by a team
B: moderateUnit trip without damage, action level 2, equipment trip with load lossApparent cause by one investigator
C: minorLimit breach caught early, action level 1, near missRecord, code and trend
  • Any repeat moves up a grade. A second minor event is no longer minor.
  • Trend the C events. Many small events in one system signal a larger one coming.
  • Decide within a day. Grade the event while the evidence is fresh.

The grades above are an illustration. Set your own criteria and write them down so grading does not depend on who is on shift.

A clear grading rule keeps investigators for the events that matter. Discuss yours with our advisors.

09Actions

Corrective Actions That Work

The strongest actions change the plant. The weakest ask people to be more careful.

1
Remove the cause

Design change: add redundancy, change the material, remove the hazard.

2
Add an engineered barrier

Interlock, alarm or automatic action.

3
Change the maintenance strategy

New task, interval or condition check.

4
Improve the procedure

Clearer steps, hold points, checklists.

5
Train and brief

Useful as support, weak on its own.

  • One action per cause. Every root and contributing cause has a matching action.
  • A named owner and date. Not a department.
  • Check other units. The same weakness may exist elsewhere.
  • Define success in advance. State what result will show the action worked.

If the only action is retraining, the investigation probably stopped too early. We review action quality in a working session.

10Follow-up

The Effectiveness Review

Checking that an action worked is a requirement in quality standards, and the step most often skipped.

ISO 9001:2015, clause 10.2
Requires organisations to determine the causes of a nonconformity, check whether similar ones exist and review the effectiveness of corrective action.
10 CFR 50 Appendix B, Criterion XVI
For US nuclear plants, requires that the cause of significant conditions is determined and “corrective action taken to preclude repetition”.
NERC EOP-004
Requires certain grid events to be reported within 24 hours, for example loss of 2,000 MW or more of generation.
When to review
After enough run time for the failure to have recurred, often three to twelve months.
What to check
Was the action done as written, and has the event or its precursors returned?
If it failed
Reopen the investigation; the root cause was not found.
100%
Actions raised
85%
Actions completed
60%
Actions verified effective

Illustrative. The gap between the second and third rings is where repeat events come from.

Scheduling the review at the moment the action is closed makes sure it happens. Try it in a pilot.

11Comparison

Paperwork-Driven vs Closed-Loop Investigation

Both produce a report. Only one changes what happens next time.

Paperwork-driven
  • Evidence gathered by hand, hours later
  • Cause recorded as the failed part
  • Report filed and forgotten
  • Actions closed when done
  • Each event stands alone
  • Same trip returns
Closed-loop
  • Event data captured automatically
  • Direct, contributing and root causes
  • Findings shared across units
  • Actions closed when proven effective
  • Events coded and trended
  • Repeat events fall
During the event
Sequence-of-events data saved
Trends and alarms captured
Statements taken separately
Event graded within a day
After the event
Causes coded to categories
One action per cause
Similar equipment checked
Effectiveness review scheduled

The checklist fits on one page beside the control desk. A process review compares it with your current practice.

12iFactory

How iFactory Closes the Loop

iFactory captures the evidence, guides the analysis and tracks every action to a verified result.

From first-out signal to proven fix, in one record.
01
Automatic event capture

Sequence of events, trends and alarms saved at every trip.

02
Guided analysis

Timeline, 5 Whys and cause coding.

03
Repeat detection

AI matches new events against history.

04
Action tracking

Owners, dates and links to work orders.

05
Effectiveness reviews

Scheduled and reported automatically.

06
Chemistry excursions

Hours at each action level, by unit.

It runs on premises and reads from your DCS, historian and laboratory data. Share your last ten trip reports and we will show the patterns in a trial.

Event review

Find the Patterns in Your Trip History

Share your trip and deviation reports for the last two years. We code the causes, match repeat events and list the actions never verified.

Illustrative alert
Unit 2 trip · 03:12

First-out was drum level low-low. Feed pump B had tripped nine seconds earlier on lube oil pressure, the same sequence as the trip in July. That corrective action was closed without an effectiveness check.

Actions verified effective60%

Window
This week
Action
Reopen the July action, inspect the lube oil system
13In practice

A Trip That Had Happened Before

This is how a shift engineer might use the system after a trip.

Shift engineer and iFactory AI
Engineer
Unit 2 tripped at 03:12. What came first?
iFactory AI
First-out was drum level low-low. Feed pump B had tripped nine seconds earlier on low lube oil pressure.
Engineer
Have we seen this before?
iFactory AI
Yes. The trip in July followed the same sequence. The action then was to replace the lube oil filter, and it was closed without an effectiveness review.
Engineer
So the filter was not the root cause.
iFactory AI
It appears not. Lube oil pressure has drifted down for three weeks since. I have reopened the July action, graded this event A as a repeat, and attached both timelines for the team.
Turnkey hardware and software

iFactory ships as a pre-configured NVIDIA AI server, racked and ready with the deviation and event investigation models loaded. Rack it, plug in power and Ethernet, and the AI is live. Scope covers data connections across units, control room, stores and planning office, DCS, historian, CMMS and ERP integration, cabling and network setup, team training and 24×7 remote monitoring.

Weeks 1–4
Ship, network, data

Server installed, DCS, historian and CMMS links live, history loaded.

Weeks 5–8
Train models, pilot

Models tuned on your own plant data, then piloted on one unit with your team reviewing every output.

Weeks 9–12
Go live, train teams

Rollout to the agreed units, team training done, 24×7 remote monitoring in place.

Software, server and integration come as one package. For pricing, contact our sales team.

FAQQuestions

Frequently Asked Questions

What is root cause analysis in a power plant?

A structured investigation that finds the cause which, if corrected, would prevent the event and similar events from recurring, then confirms the fix worked.

What are the phases of an investigation?

US Department of Energy guidance lists five: data collection, assessment, corrective actions, informing others and follow-up.

What data should be saved after a unit trip?

The sequence-of-events record, trends for at least the hour before, the alarm list, operator statements, physical evidence and the plant state at the time.

Which RCA method should we use?

For most trips, a timeline chart with 5 Whys on each causal factor. Use fault tree analysis for complex or high-consequence events.

What is an effectiveness review?

A check, some months after a corrective action, that it was carried out as written and that the event has not returned. ISO 9001 requires it.

How long does it take to set up?

Automatic event capture, cause coding and action tracking typically fit within a 6–12 week rollout. Plan it with our specialists.

Next step

Stop Investigating the Same Trip Twice

iFactory captures event evidence, guides root cause analysis and tracks every action until it is proven effective.

Illustrative dashboard view
Trips by root cause category, last 24 months
Equipment9

Design5

Procedure4

Training2

Personnel1

Illustrative. Categories follow the US DOE root cause guidance.


Share This Story, Choose Your Platform!