The reliability team at a 620 MW combined-cycle power plant did what most operations do at least once a year — they scheduled a full infrared inspection sweep, walked the plant with a handheld thermal camera, and archived the report. What they did not have was any way to see what those same panels, motors, and bearings looked like the following Tuesday, or the next month, or on the third-shift pattern when load ramped hardest. That changed 30 days after iFactory's continuous AI thermal monitoring layer went live. In the first month alone, the platform surfaced 23 previously unknown thermal anomalies — three of which the plant's own reliability engineering group later classified as forced-outage candidates. This case study walks through what the 23 findings were, how the severity framework triaged them, and the estimated dollars of unplanned outage the plant avoided. To see the same monitoring stack applied to your generation assets, book a review call with our engineering team.
23 Anomalies. 3 Forced Outages Averted. One 30-Day Window.
A 620 MW combined-cycle plant deployed iFactory AI thermal monitoring across 340 electrical panels, 68 critical motors, and 42 bearing housings. In Month 1, the system detected 23 developing hot spots that manual annual inspection had missed — three of them severe enough to have caused unplanned generation loss within 60 days.
The Plant on the Line
The site profiled here is a merchant combined-cycle power plant serving a competitive wholesale market. Availability is the number that pays the bills — every hour of forced outage costs both lost energy revenue and capacity market penalties. Manual thermal inspection was already part of their annual PM programme, but the coverage gap between inspections was exactly the window where developing faults propagate.
The 23 Findings — Categorised
Not every thermal anomaly is a forced outage waiting to happen. What matters is the mix: where the heat is showing up, how severe the deviation is against baseline, and how close the asset sits to critical generation equipment. Below is how the 23 Month-1 findings distributed across equipment class and severity band.
The Three Findings That Would Have Been Forced Outages
Every Class 4 finding was independently reviewed by the plant's reliability engineering group, who worked out the likely failure mode, the expected time-to-failure, and the estimated cost of the outage it would have caused. Each of the three below was resolved during a scheduled maintenance window at a fraction of the counterfactual cost.
Loose 480 V Bus Termination on MCC-3 Feed
A single-phase termination inside the main MCC serving cooling water pumps was running 47°C above adjacent phases under load. AI trending showed the differential had grown 8°C in the six days prior to detection. Left in place, the contact resistance would have progressed to failure — likely as an arc-flash event, taking the entire MCC offline and stopping cooling water to the steam turbine.
Boiler Feed Pump Bearing Thermal Drift
Outboard bearing housing on the 900 kW boiler feed pump motor was trending 22°C above its own six-month rolling baseline, correlated with a slight vibration rise on the same asset. Bearing degradation of this profile typically fails within 4–8 weeks. A bearing failure on the boiler feed pump means immediate loss of the HRSG water side — and a full unit trip.
Generator Excitation Panel Terminal Overheating
Two adjacent field terminals inside the excitation cubicle showed 38°C rise above baseline under full load. The AI classifier tagged the pattern as asymmetric heating consistent with a loose or oxidising contact. A failed excitation connection triggers immediate generator lock-out — with lead-time for spares running into weeks in the current market.
How the Severity Framework Actually Triages
The four-class severity scheme aligns with NETA MTS-2019 and NFPA 70B thermal anomaly guidance, adapted for continuous monitoring rather than one-shot inspection. Each detection is scored against baseline, against similar assets on the same feed, and against ambient conditions before a class is assigned and a work order is queued.
| Severity Class | Temperature Rise Over Baseline | Response Requirement | Typical Lead Time to Failure |
|---|---|---|---|
| Class 4 · Critical | >30°C above baseline or >15°C phase-to-phase | Immediate work order · Load reduction or shutdown planned within 72 hrs | 2 to 6 weeks |
| Class 3 · Serious | 15–30°C above baseline or 8–15°C phase-to-phase | Work order scheduled for next planned outage window | 1 to 3 months |
| Class 2 · Intermediate | 5–15°C above baseline · Growth trend positive | Continuous trending · Reinspect at 30-day interval | 3 to 6 months |
| Class 1 · Advisory | Elevated but stable · <5°C above baseline | Logged in condition history · No immediate action | Not imminent |
What Is Running Hot in Your Plant Right Now?
If your thermal inspection programme is annual — or even quarterly — the assets that failed silently between inspections are what will trip your next forced outage. See the continuous monitoring stack on a 30-minute walkthrough.
What Gets Monitored — and Why It Matters
Not every asset in a 620 MW plant needs continuous thermal watch. The deployment prioritised zones where thermal signatures precede failure by weeks and where failure consequence hits generation directly. This is the coverage map iFactory's reliability engineers built with the plant team during commissioning.
Loose terminations, phase imbalance, breaker contact wear, and cable termination overheating. These are the most common thermal findings in any industrial installation and the highest arc-flash risk category.
Winding insulation degradation and terminal connection heating. Motor thermal signatures typically appear 3 to 8 weeks before winding ground fault or terminal failure — enough runway for a planned rewind or replacement.
Boiler feed pumps, cooling water pumps, forced draft fans, and lube oil circulators. Bearing thermal rise correlates with lube starvation, misalignment, and rolling-element damage — leading indicators of catastrophic seizure.
Highest-consequence zone. A bushing failure or bus duct fault takes the full station offline and creates immediate life safety risk. Continuous thermal watch replaces a once-yearly inspection with per-minute observation.
Stator terminal connections, exciter rectifier assemblies, and collector ring housings. Undetected heating precedes ground faults, which in turn precede stator rewinds costing $1M and up.
Aging cable insulation develops thermal signatures as dielectric properties degrade. Early detection prevents insulation failure escalating to phase-to-phase faults and cable tray fire risk.
From Thermal Alert to Closed Work Order — Automatically
A thermal detection that sits in a report and never becomes a maintenance action is worse than no detection at all — it creates documented awareness with no documented response. iFactory closes the loop by pushing every classified anomaly through a standard workflow that ends with a signed-off work order in the plant's CMMS. This is what happened for each of the 23 findings.
Detection & Classification
AI classifier assigns a severity class based on temperature delta, growth trend, and asset criticality. Classification latency is under 90 seconds from frame capture.
Reliability Engineer Review
The on-shift reliability engineer receives a mobile alert with the thermal frame, asset ID, historical trend, and recommended action. Acknowledgement is recorded for audit.
Auto Work Order Generation
On acknowledgement, iFactory generates a work order in the plant's CMMS with the anomaly evidence, asset hierarchy, recommended trade, and severity-based priority tag.
Scheduling & Execution
Planner slots the job into the next appropriate outage window based on severity class. Class 4 goes to immediate action; Class 3 to the next planned window; Class 2 to trend and re-plan.
Post-Repair Verification
After the repair, the same thermal node re-scans the asset and confirms the anomaly is resolved. If the delta persists, the work order stays open. Closure requires thermal confirmation, not just tradesperson sign-off.
The 30-Day Financial Snapshot
Below is the cost-vs-avoidance summary the plant's asset performance team presented at Month 1 review. Numbers reflect actuals for the repairs executed and reliability-engineering estimates for the outages that did not happen.
| Line Item | Amount | Category |
|---|---|---|
| Planned repairs for 3 Class 4 findings | $32,400 | Incurred |
| Planned repairs for 7 Class 3 findings | $54,700 | Incurred |
| Trend monitoring for 9 Class 2 findings | $0 | No spend |
| Advisory logging for 4 Class 1 findings | $0 | No spend |
| Estimated avoided cost · Forced outage 1 (MCC bus) | $1,800,000 | Avoided |
| Estimated avoided cost · Forced outage 2 (BFP bearing) | $2,400,000 | Avoided |
| Estimated avoided cost · Forced outage 3 (Excitation panel) | $3,100,000 | Avoided |
| Net 30-Day Value | $7,212,900 | Return |
Avoided outage costs are estimated using the plant's own forced-outage cost model at prevailing energy prices during the assessment window. Actual outcomes depend on failure mode timing and market conditions at the point of failure.
Frequently Asked Questions
How is a thermal anomaly different from a normal hot spot on operating equipment?
Every energised asset produces heat under load — that is not an anomaly. A thermal anomaly is a deviation from what that specific asset should look like given its baseline, its load, and the ambient conditions at the moment of measurement. iFactory establishes a rolling baseline for each monitored asset from 60 to 90 days of continuous data and only alerts when the current thermal signature exceeds that baseline by a defined threshold or shows a positive growth trend. This asset-specific approach avoids the false alarms that generic temperature-limit alerts produce. To review how baselining works on your equipment mix, book a session with our engineers.
Can AI thermal monitoring run on a plant network with strict cybersecurity requirements?
Yes. The platform is designed for deployment inside the plant control network with zero cloud dependency. All inference, storage, and dashboarding run on an on-premise appliance that meets NERC CIP and, where required, nuclear security control room requirements. Alerts and work orders are pushed to authorised endpoints inside the same network segment; no thermal data or asset metadata leaves the facility unless explicitly configured for corporate rollup. For a detailed architecture review with your OT security team, reach our support team.
Does the system replace annual thermographic inspections required by insurance?
In most cases the answer is a productive negotiation with your insurer rather than an outright replacement. Continuous thermal monitoring produces a far more complete evidentiary record than a once-annual sweep, and several insurers now allow reduced inspection cadence or premium adjustments when a documented continuous monitoring programme is in place. iFactory produces exportable, timestamped anomaly logs and NETA-aligned severity classifications specifically to support insurance and audit conversations. To see an example insurer-format export, book a walkthrough.
What happens when equipment is added, moved, or replaced after go-live?
Every asset in the monitoring scope carries its own model configuration and baseline. When new equipment is added or existing equipment is replaced, the reliability engineer initiates an asset-onboarding task that captures a new baseline over the appropriate learning window — typically 30 to 60 days for stable operating equipment. Until the new baseline is established, the asset runs on similar-asset benchmarks so it is never unmonitored. Change control documentation is generated automatically to keep the audit trail intact. For questions about specific asset transitions on your site, contact our team.
How is this different from just installing more fixed thermocouples?
Fixed thermocouples read one point per sensor. AI thermal imaging reads the full surface of the asset — thousands of temperature points per frame — and it does so through infrared imaging that does not require physical contact with energised equipment. A loose termination that heats the connection lug 30°C hotter than the surrounding bus will be caught by thermal imaging; a thermocouple mounted 100 mm away will not show the delta. Fixed sensors and thermal imaging are complementary, not substitutable — most sites run both, with thermal imaging providing the wide-area early warning and thermocouples confirming trip-level readings. To see how the two integrate, schedule a demo.
Turn Month 1 Into Your Own Result
Every plant we deploy on finds a Class 4 anomaly within the first 60 days. The question is whether you find it before it trips your unit, or after. Let us map coverage for your critical assets and show you what continuous thermal monitoring would surface.







