Every cold-rolled sheet line running above 1,200 meters per minute generates more surface area per hour than a human inspector can meaningfully assess in a week. QA Vision Leads who have deployed AI surface inspection systems know that the real challenge is not getting a system to detect defects — it is getting the system to classify them correctly. A model that flags 98 percent of surface anomalies but misclassifies scratches as scuffs, or oil stains as edge dents, produces data that downstream graders, metallurgists, and process engineers cannot act on. The defect taxonomy — the specific classes the model recognizes, the boundaries between those classes, and the published accuracy and confusion rates for each — is what separates a system that earns trust from one that gets ignored. iFactory's defect taxonomy is built around this principle: published per-class accuracy, transparent confusion matrices, and retrainable models that adapt to your specific mill configuration.
Why the Taxonomy Problem Is Harder Than the Detection Problem
Binary defect detection — is there something on this surface or not — is a solvable problem with mature convolutional neural network architectures. The classification problem — what exactly is that something — is where AI vision systems diverge dramatically in their real-world value. Consider two defects that are visually adjacent: a scratch, which is a linear groove cut into the surface by a foreign object or a damaged roll, and a scuff, which is a superficial abrasion that displaces surface material without cutting into the substrate. On a moving strip at 1,200 meters per minute under variable lighting, the pixel-level difference between a scratch and a scuff can be less than two percent of the total defect pixel area. A model that has not been specifically trained to respect that boundary will classify them interchangeably, which means your scratch trend data is contaminated with scuff events and vice versa.
For a QA Vision Lead, this is not an academic distinction. Scratch root causes — roll damage, strip guide misalignment, debris in the roll bite — are completely different from scuff root causes — excessive inter-stand tension, roll surface finish degradation, improper coil handling. If the classification is unreliable, the root cause analysis that depends on it is unreliable, and the process engineers stop trusting the system's output within weeks. The defect taxonomy must be designed from the outset around class boundaries that map to distinct root causes and distinct corrective actions, not just distinct visual appearances.
The six-class taxonomy below represents the defect categories that iFactory has validated across multiple cold-rolling installations, with per-class accuracy rates measured against human-expert-labeled datasets of more than 50,000 defect instances per class. Each class is defined by its physical characteristics, its root cause category, and the specific imaging conditions required for reliable detection.
Full Defect Class Map — Cold-Rolled Sheet
Each class below includes the baseline accuracy rate achieved on the iFactory reference dataset, the typical false-positive rate observed across deployed installations, and the root cause category that makes the class actionable for process engineers. The accuracy bars represent classification precision — when the model says a defect is class X, how often it is actually class X.
Localized deformation at the strip edge caused by coil handling, improper uncoiler tension, or impact from guide rollers. Typically appears as a smooth, rounded indentation 5 to 30 millimeters in length along the edge 10 to 50 millimeters from the strip boundary. Detected using angled edge lighting that creates shadow contrast in the deformation zone. Root cause category: mechanical handling and tension control.
Residual rolling oil or emulsion that is not uniformly distributed across the strip surface, appearing as irregularly shaped dark patches with diffuse boundaries. Severity ranges from light film variations that evaporate during downstream annealing to heavy pooling that causes carbon deposition and surface discoloration. Detected using diffuse axial lighting at specific wavelengths where oil absorbs differently than bare metal. Root cause category: lubrication system and emulsion spray configuration.
Linear groove cut into the strip surface by a foreign object, a damaged work roll, or a misaligned component in the strip path. Characterized by sharp edges, consistent width along the length, and depth that penetrates below the surface oxide layer. Length ranges from a few millimeters to the full strip length for continuous roll-contact scratches. Detected using directional lighting aligned perpendicular to the expected scratch orientation. Root cause category: roll condition, strip path alignment, and debris exclusion.
Superficial abrasion that displaces surface material without cutting into the substrate, typically caused by excessive friction between the strip and roll surfaces or between the strip and guide components. Visually similar to scratches but with diffuse edges, variable width, and no consistent depth profile. This is the hardest class in the taxonomy due to its visual overlap with scratches and its sensitivity to lighting angle. Root cause category: inter-stand tension, roll surface finish degradation, and guide contact pressure.
Characteristic cross-hatched pattern of surface ridges oriented at approximately 45 degrees to the rolling direction, caused by work roll chatter vibration or by strip slip between stands during speed transitions. The pattern is highly regular and visually distinctive, which makes it the easiest class for AI to identify with high precision. Severity is graded by ridge height and the percentage of strip width affected. Root cause category: mill vibration, chatter control, and inter-stand speed synchronization.
Fractures initiating at the strip edge and propagating inward, caused by excessive reduction per pass, insufficient edge trimming, or material with marginal edge ductility. Cracks range from hairline fractures visible only under magnification to open splits several centimeters long. Detected using edge-specific camera stations with backlighting that creates high contrast at the crack boundary. Root cause category: reduction schedule, edge trim configuration, and incoming material edge quality.
The Confusion Matrix — Where Classification Breaks Down
A headline accuracy number like 93 percent tells you nothing about where the model fails. The confusion matrix below shows exactly how each defect class is misclassified — which classes get confused with which other classes, and at what rate. This is the document a QA Vision Lead needs to evaluate whether a model is ready for production use, because it reveals the specific confusion pairs that will contaminate your defect trend data if left unaddressed.
| Edge Dent | Oil Stain | Scratch | Scuff | Herringbone | Edge Crack | |
|---|---|---|---|---|---|---|
| Edge Dent | 94.2 | 0.8 | 1.2 | 0.5 | 0.3 | 3.0 |
| Oil Stain | 0.5 | 91.3 | 0.3 | 2.8 | 0.4 | 4.7 |
| Scratch | 0.4 | 0.2 | 93.1 | 5.2 | 0.3 | 0.8 |
| Scuff | 0.6 | 1.1 | 8.3 | 86.7 | 0.9 | 2.4 |
| Herringbone | 0.2 | 0.1 | 0.4 | 0.6 | 97.8 | 0.9 |
| Edge Crack | 2.1 | 0.9 | 0.7 | 1.8 | 0.5 | 94.0 |
What Drives the Three Critical Confusion Pairs
Three confusion pairs in the matrix above account for the vast majority of classification errors, and understanding the physical basis for each pair is essential for deciding whether to accept the current performance or invest in plant-specific retraining. Each pair below explains why the model confuses the two classes and what plant-specific factors can reduce the confusion rate.
This is the single largest confusion pair in the taxonomy. At line speed, shallow scratches — where the groove depth is less than 5 microns — produce surface reflections that are nearly indistinguishable from scuff marks under standard lighting. The model relies on edge sharpness and width consistency along the defect length to distinguish the two classes, but when the scratch is shallow and the strip surface has minor roughness variation, those features become unreliable. Plant-specific retraining with shallow scratch samples from your specific roll grades and strip alloys typically reduces this confusion from 5.2 percent to below 3 percent by teaching the model the subtle depth cues that are unique to your surface finish range.
The reverse confusion is even larger because scuff marks are inherently more variable in appearance than scratches. A scuff that occurs at a guide roller where the contact pressure is uneven can produce a linear abrasion pattern that mimics a scratch in both shape and orientation. The model classifies these as scratches because the linear morphology matches the scratch template more closely than the diffuse scuff template. This confusion is particularly prevalent on lines where guide roller surfaces are worn or where strip tension fluctuates during speed changes, because both conditions produce scuff patterns that look more linear than the training data expects. Adding scuff samples from your specific guide roller configurations to the training set is the most effective correction.
This confusion occurs when oil pooling concentrates near the strip edge and creates a dark, irregular boundary that the model interprets as a crack opening. The spectral signature of pooled oil and the shadow pattern at a crack edge share enough low-frequency features to cause misclassification when the oil stain is large and the edge camera's lighting angle creates partial shadow effects at the oil boundary. Plants that run heavy emulsion applications or that have marginal edge wiping systems see this confusion pair more frequently. The correction is a combination of adding edge-region oil stain samples to the training set and adjusting the edge camera lighting angle to reduce shadow overlap with the oil boundary.
Per-Class Detection Complexity and False-Positive Drivers
False positives are more damaging to system adoption than false negatives in most cold-rolling environments, because a false positive triggers unnecessary downgrading or reinspection that costs real money, while a false negative on a minor defect may be caught by downstream inspection. The cards below break down what specifically drives false positives for each class, so QA teams can assess whether their plant's operating conditions are likely to produce higher or lower FP rates than the baseline.
Edge waviness caused by strip shape control issues creates shadow patterns under angled lighting that resemble dent profiles. The model can distinguish a true dent from waviness when the shadow gradient is steep enough, but mild waviness with gradual depth transitions falls within the dent feature space. Plants with ongoing shape control issues will see FP rates above the 1.8 percent baseline. Adjusting the edge lighting angle to reduce waviness shadow intensity without losing dent contrast is the primary mitigation.
Normal emulsion drainage patterns on the strip surface create variations in reflectivity that can trigger the oil stain class when the pattern is locally concentrated. Water spots from imperfect drying also produce dark patches that overlap with the oil stain feature space. The FP rate is highly sensitive to the emulsion concentration and the drying system performance. Plants running high emulsion ratios or with aging dryer nozzles typically see FP rates of 4 to 6 percent on this class until the model is retrained with samples that include normal emulsion variation patterns specific to that plant.
Periodic roll marks and light chatter lines produce linear features on the strip surface that share morphological characteristics with scratches. The model distinguishes scratches from roll marks primarily by periodicity — roll marks repeat at the roll circumference interval, while scratches are typically singular or randomly spaced. When the roll mark periodicity is irregular due to roll eccentricity or when a single roll mark is isolated within the camera field of view, the periodicity cue is lost and the model classifies it as a scratch. Including plant-specific roll mark samples with known periodicity characteristics in the training set resolves most of these false positives.
Cold-rolled strip surface roughness varies across the width due to roll crown profiles and backup roll flattening patterns. Areas of naturally higher roughness scatter light differently and can produce a mottled appearance that triggers the scuff class. Temporary coating applied for corrosion protection during storage also creates surface texture variations that overlap with scuff features. The scuff class has the highest baseline FP rate at 5.6 percent precisely because it is defined by the absence of sharp features rather than the presence of distinctive ones. Retraining with roughness map data from your specific roll campaigns is the most effective reduction strategy.
The cross-hatched geometry of herringbone defects is so visually distinctive that false positives on this class are rare at 0.4 percent. The few false positives that do occur are typically caused by directional surface finishing marks from backup rolls that happen to approximate the herringbone angle, or by light interference patterns from overlapping lighting sources. This class is the safest starting point for validating a new installation because any detections are almost certainly real, which builds operator confidence in the system before moving to the harder classes.
Edge trimming and slitting operations leave micro-burr and shear marks along the strip edge that create discontinuities in the edge profile detected by the edge camera. When the burr height is above 10 microns and the lighting angle creates a shadow on the burr side, the edge profile shows a discontinuity that the model interprets as a crack initiation point. Plants with dull slitter knives or excessive trimmer knife clearance see elevated FP rates on this class. Retraining with edge profiles that include burr patterns at various knife wear stages, combined with a knife change schedule that keeps burr height below the detection threshold, brings the FP rate back to baseline.
Plant-Specific Retraining — Why Generic Models Fail on Your Line
The baseline accuracy and confusion matrix numbers published above are measured on iFactory's reference dataset, which includes defect samples from multiple cold-rolling installations across different steel grades, roll configurations, and lighting setups. That dataset is deliberately broad to ensure the model generalizes well across the industry. But no generic dataset captures the specific combination of surface finish characteristics, lighting geometry, defect morphology variations, and false-positive trigger conditions that exist on your specific line. Plants that deploy the generic model without retraining typically see overall accuracy 4 to 7 percentage points below the published baseline, with the gap concentrated in the hardest classes — scuff and oil stain — where plant-specific variation matters most.
The retraining process requires a minimum of 2,000 labeled samples per class from your specific line to produce a statistically meaningful improvement. Those samples must include not only the defect instances but also the false-positive trigger conditions — edge waviness, emulsion patterns, roll marks, burr profiles — that are specific to your plant. The expert labeling step is where most plants underestimate the effort: every sample must be reviewed and classified by a senior QA engineer who understands the physical distinction between classes, not just the visual one. A sample that looks like a scuff to a junior inspector may be correctly identified as a shallow scratch by an engineer who knows the roll condition history at the time the sample was collected. That domain expertise is what turns a retrained model from slightly better to significantly better.
Before and After Plant-Specific Retraining
The table below shows representative accuracy improvement from plant-specific retraining at a 5-stand tandem cold mill processing 0.4 to 2.0 mm strip in low-carbon and IF steel grades. The before column reflects the generic model performance in the first two weeks of shadow-mode operation. The after column reflects performance after retraining with 3,200 plant-specific samples per class and four weeks of additional validation. The improvement is not uniform — it concentrates in exactly the classes where plant-specific variation drives false positives.
| Defect Class | Generic Model Precision | Retrained Precision | Improvement | FP Rate Before | FP Rate After |
|---|---|---|---|---|---|
| Edge Dent | 89.4% | 95.1% | +5.7 pts | 4.2% | 1.4% |
| Oil Stain | 84.7% | 93.6% | +8.9 pts | 7.1% | 2.8% |
| Scratch | 88.9% | 94.8% | +5.9 pts | 4.8% | 1.6% |
| Scuff | 79.3% | 89.4% | +10.1 pts | 9.8% | 4.2% |
| Herringbone | 96.1% | 98.2% | +2.1 pts | 0.9% | 0.3% |
| Edge Cracking | 90.2% | 95.3% | +5.1 pts | 3.9% | 1.7% |
A model with 94 percent overall accuracy and a 6 percent false-positive rate will be disabled by operators within three months. A model with 89 percent overall accuracy and a 1.5 percent false-positive rate will be adopted and gradually improved. The reason is purely operational: every false positive requires someone to walk to the line, inspect the flagged location, determine that the defect is not real, and clear the alarm. At high line speeds with frequent false positives, this activity consumes more operator time than the manual inspection it was supposed to replace. QA Vision Leads who have lived through failed AI deployments consistently identify false-positive fatigue as the number one reason operators stop trusting the system. The published per-class false-positive rates in iFactory's defect taxonomy exist specifically so that you can evaluate the FP exposure before deployment, set realistic expectations with your operations team, and prioritize the retraining effort on the classes where FP reduction will have the largest operational impact.






