A plant reports that OEE improved four points in the same quarter an AI-driven maintenance program went live, and a skeptical CFO asks the one question that most impact reports can't actually answer: how do you know it was the AI and not the new shift supervisor, the seasonal product mix shift, or simple regression to the mean after an unusually bad prior quarter? Correlation between an AI rollout and an improved metric is not the same as attribution, and the gap between the two is exactly where AI programs lose credibility with finance even when the underlying impact is real. Rigorous measurement requires a defined counterfactual, explicit attribution rules agreed before the program launches, and a reporting cadence that shows the trend holding over time rather than a single favorable snapshot. If your last impact report got a raised eyebrow instead of a nod, book a demo to see a measurement framework that survives CFO scrutiny.
"How Do You Know It Was the AI?" Should Have a Rehearsed Answer.
Correlation between an AI rollout and an improved metric isn't proof. iFactory's measurement framework builds a defined counterfactual, explicit attribution rules, and a reporting cadence that holds up when a skeptical CFO asks the obvious question.
Why "The Metric Improved After We Launched AI" Isn't Proof
FMCG plants have dozens of variables moving simultaneously — seasonal demand shifts, staffing changes, raw material quality variation, equipment age, and a running list of other process improvements happening in parallel. Attributing a metric improvement to a specific AI program without accounting for these confounding factors is a claim finance is right to question.
Without a comparison group or a modeled baseline of what would have happened without the AI program, there's no way to isolate its specific contribution from everything else changing at once.
A metric that naturally improves during a specific season can be mistaken for AI impact if the comparison period isn't adjusted for known seasonal patterns.
When an AI program launches alongside other operational improvements, isolating which initiative drove which portion of the result requires explicit attribution logic agreed in advance.
A single favorable data point right after launch can reflect short-term novelty effects or regression to the mean rather than a durable, repeatable improvement.
Building a Counterfactual That Actually Holds Up
A counterfactual answers the specific question "what would have happened without this program," and there are several established approaches to constructing one credibly in a plant environment.
Matched Control Lines
Comparing a line running the AI program against a similar line running the prior process, matched as closely as possible on product mix, age, and staffing, isolates the AI's specific contribution.
Trailing Baseline Trend
Extending the pre-launch trend line forward as the expected trajectory without intervention, then measuring the actual result against that projected trend rather than a flat prior average.
Staggered Rollout Comparison
Rolling out to different lines or plants at different times creates a natural comparison group of not-yet-live sites to benchmark against during the rollout period.
Design a Measurement Framework Before You Launch, Not After
See how to set up counterfactual comparisons and attribution rules before your next AI program goes live, so the impact report is credible from day one.
Attribution Approach by Metric: OEE, Inventory, and Waste
Each of these three commonly claimed metrics has its own specific measurement challenges and requires a slightly different attribution approach to isolate AI-driven impact credibly.
| Metric | Key Measurement Challenge | Recommended Attribution Approach |
|---|---|---|
| OEE | Multiple concurrent improvement initiatives affecting the same lines | Matched control line comparison, isolating downtime specifically tied to AI-flagged interventions |
| Inventory | Seasonal demand variation and promotional activity affecting stock levels | Trailing baseline trend adjusted for known seasonal and promotional calendar effects |
| Waste | Raw material quality variation independent of any process change | Waste category breakdown isolating the specific defect types the AI program targets |
A Reporting Cadence That Builds Credibility Over Time
A single impact number reported once, right after a favorable quarter, invites exactly the skepticism a rigorous cadence is designed to avoid. Showing the trend holding over multiple reporting periods is what actually convinces a skeptical audience.
Establish the counterfactual baseline and confirm data collection methodology is consistent across the comparison group before any impact claim is made.
First directional read against the counterfactual, reported with appropriate caution and explicitly labeled as early and provisional rather than final.
First credible impact claim, if the trend has held consistently across the full period rather than showing a single favorable spike.
Continued tracking against the counterfactual to confirm the impact is durable rather than a temporary launch effect that fades over time.
The Impact Report That Survived Board Scrutiny
A dairy processor's operations team reported a 6-point OEE improvement on the line running a new AI-driven predictive maintenance program, presenting the number confidently to the executive team as clear evidence of success. The CFO's first question — "what would OEE have done without this program" — had no prepared answer, and the credibility of the entire initiative took a visible hit in that meeting despite the underlying result likely being genuine.
For the next reporting cycle, the team built a matched control line comparison using a similar line at a sister plant that hadn't yet received the AI program, tracked both lines' OEE over a full two-quarter period, and adjusted for a known product mix change that affected both lines equally during the period. The revised report showed the AI-equipped line improving 5.2 points relative to the control line's own 0.8-point improvement over the same period — a more modest but far more defensible number that the CFO accepted without further challenge, and that became the template for how every subsequent AI program's impact was measured and reported.
Frequently Asked Questions
What if we don't have a comparable control line or plant to use as a counterfactual?
A trailing baseline trend is a reasonable alternative when a true matched control isn't available — extending the pre-launch trend line forward as the expected trajectory and comparing actual post-launch performance against that projection rather than a flat historical average. This approach is less rigorous than a genuine control group comparison but still meaningfully more credible than a simple before-and-after comparison with no adjustment for the underlying trend the metric was already following. Staggered rollout timing, if your deployment plan allows for it, is another option worth considering even after the fact for future programs. Our team can help design the right counterfactual approach for your specific plant configuration — book a demo to review your options.
How do we handle a situation where multiple improvement initiatives launched around the same time as the AI program?
This requires explicit attribution rules agreed before analyzing results, ideally isolating the specific mechanism by which the AI program is expected to drive improvement — for OEE, this might mean tracking only downtime events specifically flagged by the AI-driven maintenance recommendations, rather than claiming credit for the full OEE change when other initiatives were also targeting different downtime categories simultaneously. Where full isolation isn't possible, transparently acknowledging the overlap and presenting a range rather than a single precise attributed figure maintains credibility better than an overstated precise claim. For guidance on structuring attribution rules for concurrent initiatives, contact our support team.
How long should we wait before reporting any impact number to leadership?
Waiting until a full six-month trend is established before making a formal impact claim is a reasonable standard for most FMCG metrics, since this window is generally long enough to smooth out short-term novelty effects and seasonal noise while still being timely enough to inform ongoing investment decisions. Earlier directional updates are appropriate and expected, but should be explicitly labeled as provisional rather than presented with the same confidence as a fully validated result. Reporting too early with an unvalidated number risks the exact credibility damage a rigorous framework is meant to avoid. For a specific reporting timeline tailored to your metric and program, schedule a session with our team.
Does this level of measurement rigor apply to smaller, lower-investment AI initiatives too?
The level of rigor should scale with the size and visibility of the claim being made — a small pilot with a modest investment doesn't need the same formal control-line methodology as a multi-million dollar enterprise rollout, but even a small initiative benefits from at minimum a trailing baseline comparison rather than a bare before-and-after snapshot. The core principle — some form of counterfactual, however simple, before claiming attribution — scales down reasonably well even for lightweight pilots. For a proportionate measurement approach for a smaller initiative, reach out to support.
Who should own building and presenting the impact measurement, the operations team or finance?
A collaborative approach produces the most credible result — operations typically has the domain knowledge to identify confounding factors and design an appropriate counterfactual, while finance brings independent verification credibility that meaningfully increases the audience's trust in the final number, particularly when the claim is being presented to a board or executive committee. A number independently validated or co-produced with finance carries more weight than the same number presented solely by the team whose program is being evaluated. For guidance on structuring this collaboration, book a demo to discuss your specific reporting structure.
Build an Impact Report That Doesn't Fold Under the First Hard Question
See how iFactory structures counterfactual comparisons and attribution rules so your OEE, inventory, and waste claims hold up to real scrutiny.







