Steel AI Pilot: A 12-Week Implementation Roadmap

By James Smith on July 20, 2026

steel-ai-pilot-12-week-implementation

Most steel AI pilots die in the demo zone. The vendor arrives with a polished slide deck, a working demo on a competitor's data, and a promise to "show value in 90 days." Three months in, the plant has a dashboard nobody logs into, a model that flags every casting event as an anomaly, and a project manager walking into a Phase 2 review with no defensible business case. The problem is almost never the model — it is that the pilot was scoped to prove the vendor's product works, not to answer the questions your CFO will ask. Book a demo to walk through a 12-week framework built to survive that review.

PROJECT MANAGER · STEEL · 12-WEEK PILOT ROADMAP

A 12-Week Steel AI Pilot Framework That Survives the Phase 2 Funding Review — Not Just the Demo

The pilots that scale share four traits: a narrow scoped use case, a defensible baseline captured before the model is turned on, four explicit gate criteria, and a Phase 2 business case drafted in week one — not week twelve.

12 Wk
Standard Pilot Window From Kickoff to Gate Review
4
Explicit Phase Gates Between Pilot and Phase 2
40-70%
AI Pilots Reported to Stall Before Scaling
WHY PILOTS STALL

The Four Reasons Steel AI Pilots Die in the Demo Zone

Every stalled steel AI pilot post-mortem lands on the same four themes. None are about model accuracy — they are about how the pilot was scoped, sequenced, and measured before the model was trained. A project manager who catches these at kickoff turns 12 weeks into a defensible Phase 2 request. Missing them turns the pilot into a science project the CFO defunds.

R1
Scope Set to "Show What AI Can Do" Instead of a Specific Failure Mode
Pilots that try to demonstrate the whole platform never land a testable outcome. The pilots that scale start with one narrow failure mode — caster breakout, mill bearing degradation, reheat efficiency — and prove or disprove value against that target.
R2
Baseline Never Captured Before the Model Was Deployed
Without a documented pre-model baseline, there is no defensible way to prove AI improved anything. Six months later the plant manager points to a good quarter and says the model made no difference — and is right, because nobody wrote down the starting point.
R3
Success Criteria Defined Only After the Data Is In
If the target metric is chosen at week eleven based on which chart looks best, the pilot has lost credibility. Gate criteria are agreed in week one and signed by the sponsor before any tag is wired — that is what makes the Phase 2 case defensible to finance.
R4
No Business Case Drafted Until the Pilot Is Already Over
A Phase 2 business case built in week twelve is a scramble. The pilots that scale draft the ROI model in week one — savings per prevented event, cost per false positive, capex for full rollout — and refine the numbers as the pilot produces actuals against the model.
THE 12-WEEK PILOT FRAMEWORK

Four Phases, Four Gate Reviews, One Defensible Phase 2 Case

The 12 weeks split into four three-week phases, each ending in a gate review with the executive sponsor. Pass the gate, phase continues. Miss the gate, pilot pauses for corrective action — not silently drifts into month six.

Weeks 1-3
Scope & Baseline
  • Use case narrowed to one failure mode with one asset class
  • Historical baseline captured over the last 12 months
  • Four gate criteria signed by the executive sponsor
  • Draft Phase 2 business case circulated in week two
Gate 1 · Scope & Baseline Signed
Weeks 4-6
Data Wiring & Model Setup
  • Historian, MES, and SAP read paths provisioned and tested
  • Model trained on 12-24 months of plant-specific history
  • Alert taxonomy agreed with shift leads and maintenance
  • Shadow mode readiness confirmed for the target zone
Gate 2 · Data & Model Ready
Weeks 7-9
Shadow Mode & Tuning
  • Model runs in read-only shadow against live plant data
  • Every alert reviewed by a qualified engineer or operator
  • False positive rate tracked and tuned each week
  • Accept-or-override log captured for Phase 2 evidence
Gate 3 · Shadow Mode Passed
Weeks 10-12
Business Case & Handover
  • Documented events flagged and their business outcomes traced
  • ROI model updated with actual pilot data, not vendor claims
  • Phase 2 scope, timeline, and capex drafted for sponsor review
  • Executive handover deck ready for the funding committee
Gate 4 · Phase 2 Case Approved
GATE CRITERIA SCORECARD

The Success Criteria That Turn a Steel AI Pilot Into a Fundable Phase 2 Case

Every phase gate has explicit pass criteria agreed in week one. If any criterion misses, the phase does not close silently — the sponsor reviews and either accepts a remediation plan or ends the pilot. The scorecard below is what a good pilot walks into every gate meeting with.

Gate Pass Criterion Evidence Required
Gate 1 Scope narrowed & baseline documented Signed one-page scope + 12-month baseline data
Gate 2 Data wired & model trained Read-path test log + model performance on holdout
Gate 3 Shadow mode meets alert quality target False positive rate + operator accept-rate log
Gate 4 Phase 2 case defensible to finance ROI model with pilot actuals + Phase 2 capex plan
Any Gate Executive sponsor present at review Attendance confirmed + decision minute captured

The Pilots That Scale Look Boring on the Kickoff Slide — and Bulletproof at the Funding Review

Narrow scope, documented baseline, four gate criteria signed by the sponsor, and a Phase 2 case drafted in week one — the framework the pilots that survive share.

ROLES ON THE PILOT TEAM

Who Actually Has to Be in the Room for a Steel AI Pilot to Succeed

A steel AI pilot with only IT and vendor staff in the room is already halfway to stalling. The pilots that scale bring together a specific set of roles from week one, each with explicit decision authority. If any of the six roles below is missing at kickoff, the project manager should escalate before spending on data wiring.

Executive Sponsor
Owns the Phase 2 funding decision. Attends every gate review. Signs scope and gate criteria in week one — without this signature, no pilot should start.
Plant Manager or Ops Director
Owns the operational outcome the pilot targets. Approves shadow-mode exposure and confirms the target zone is available for the pilot window.
Maintenance or Process Lead
Subject-matter expert on the failure mode. Reviews every model alert during shadow mode and provides the accept-or-override signal that tunes the model.
IT / OT Integration Lead
Owns historian, MES, and SAP read paths. Ensures data wiring in weeks 4-6 respects ISA-95 boundaries and network segmentation.
Finance Partner
Signs off the ROI model in week one and updates it with pilot actuals. Represents the funding committee's questions before they hit Phase 2 review.
Project Manager
Runs the 12-week cadence. Manages gate reviews, tracks evidence, escalates missed criteria within the phase, and delivers the executive handover deck.
SCALING SIGNALS

The Signals a Pilot Is Actually Ready to Scale Into Phase 2 Funding

Not every pilot that finishes 12 weeks should scale. The decision to promote into Phase 2 rests on four signals — not on optimism about the technology or momentum from the demo. A project manager who can point to all four walks into the funding review with a defensible case. Missing any one is a signal to run a second iteration.

S1
Documented Alert-to-Outcome Trace
At least three shadow-mode alerts that map to an actual outcome — a prevented breakout, a caught bearing wear event, a defect trend surfaced ahead of downgrade. Vague claims do not survive finance scrutiny.
S2
False Positive Rate Trending Down
Weekly false positive rate falls across the shadow window. A flat curve means the model has not learned the plant. A rising curve means the target is wrong.
S3
Operator Accept Rate Above Threshold
Qualified operators accept model recommendations above the agreed threshold — typically 60-70% in shadow mode. Below that, either the model is not ready or the alert taxonomy needs redesign.
S4
Sponsor Willing to Sign the Phase 2 Scope
Ultimately the pilot succeeds when the sponsor signs the Phase 2 scope. If the sponsor hesitates at gate 4, the project manager needs a second pilot iteration before requesting capex — not a bigger deck.
FREQUENTLY ASKED QUESTIONS

Steel Project Managers' Questions About the 12-Week AI Pilot Framework

Can a steel AI pilot be compressed into less than 12 weeks if the executive sponsor wants faster proof?
In principle yes, in practice rarely well. The 12-week structure exists because each phase has a real gating item — scope agreement, data wiring, shadow-mode learning, business case. Compressing weeks 1-3 typically means the baseline is not captured and gate 4 has no defensible comparison. Compressing weeks 4-6 usually means data quality issues surface in week eight. If speed matters most, narrow the scope further rather than shrink the timeline. Book a scoping session to structure a compressed but defensible pilot.
What if our steel plant does not have 12-24 months of clean historical data to train the model on?
This is one of the most common realities in mid-sized steel plants and does not disqualify a pilot. The framework adjusts: the model uses whatever historical data is available for initial calibration, and gate 3 shadow-mode performance carries more weight in the Phase 2 case than model training quality. Some failure modes — bearing degradation, thermal drift — calibrate faster in shadow mode than on historical data, because the plant-specific signature emerges in weeks 7-9. Talk to pilot support to review the data available in your plant.
Should we run the pilot with our incumbent automation vendor or bring in a specialist AI platform for the 12 weeks?
This is the classic build-versus-buy question and the honest answer depends on what the incumbent vendor can demonstrate at kickoff. If they have production references in steel with a comparable failure mode, running the pilot with them lowers integration risk. If their AI product is a slide-deck extension of their DCS or MES, a specialist platform is usually the safer 12-week bet. Either way, the framework — narrow scope, four gate criteria, shadow mode — stays the same. Book a session to review vendor options against your pilot scope.
How do we handle the situation where the pilot succeeds technically but the plant team is not ready to change how they work?
This is the outcome that Phase 2 funding reviews most often kill, and it is why the maintenance or process lead is in the room from week one. Change readiness is a gate 4 criterion in every well-run pilot — not an afterthought. If the plant team sees the model as a threat rather than a tool by week nine, the project manager reworks the alert taxonomy and shadow workflow before requesting scale funding. Contact pilot support to review change-readiness practices for steel.
What does a defensible Phase 2 capex request from a steel AI pilot typically look like at the funding review?
A defensible Phase 2 request has three elements: pilot evidence (alert-to-outcome trace, false positive trend, operator accept rate), scope for rollout (which additional assets, zones, or plants and in what sequence), and a capex plan tied to timelines. What it does not include is a projection based on vendor benchmarks alone. Finance committees fund what they can trace back to the pilot's own data — that is the entire point of the 12-week framework. Book a session to structure a Phase 2 request against your pilot outputs.
TURN YOUR NEXT PILOT INTO A FUNDABLE PHASE 2

Give Your Next Steel AI Pilot the Structure the Funding Committee Is Actually Looking For

Narrow scope, documented baseline, four gate reviews, and a Phase 2 case drafted in week one — that is the framework the pilots that scale share. Book a session to map the 12-week pilot onto your steel operation.


Share This Story, Choose Your Platform!