AI Vision Defect Classes for Cold-Rolled Sheet — Full Class Map

By Johnson on July 25, 2026

ai-vision-defect-classes-cold-rolled-sheet

Every cold-rolled sheet line running above 1,200 meters per minute generates more surface area per hour than a human inspector can meaningfully assess in a week. QA Vision Leads who have deployed AI surface inspection systems know that the real challenge is not getting a system to detect defects — it is getting the system to classify them correctly. A model that flags 98 percent of surface anomalies but misclassifies scratches as scuffs, or oil stains as edge dents, produces data that downstream graders, metallurgists, and process engineers cannot act on. The defect taxonomy — the specific classes the model recognizes, the boundaries between those classes, and the published accuracy and confusion rates for each — is what separates a system that earns trust from one that gets ignored. iFactory's defect taxonomy is built around this principle: published per-class accuracy, transparent confusion matrices, and retrainable models that adapt to your specific mill configuration.

AI VISION DEFECT CLASSES · COLD-ROLLED SHEET · CONFUSION MATRIX PUBLISHED
The Defect Class Map That Determines Whether Your AI Gets Trusted or Silenced
Edge dent, oil stain, scratch, scuff, herringbone, edge cracking — six defect classes that account for the vast majority of cold-rolled sheet rejections. iFactory publishes per-class accuracy, false-positive rates, and the full confusion matrix so your QA team knows exactly where the model is reliable and where it needs plant-specific retraining.

Why the Taxonomy Problem Is Harder Than the Detection Problem

Binary defect detection — is there something on this surface or not — is a solvable problem with mature convolutional neural network architectures. The classification problem — what exactly is that something — is where AI vision systems diverge dramatically in their real-world value. Consider two defects that are visually adjacent: a scratch, which is a linear groove cut into the surface by a foreign object or a damaged roll, and a scuff, which is a superficial abrasion that displaces surface material without cutting into the substrate. On a moving strip at 1,200 meters per minute under variable lighting, the pixel-level difference between a scratch and a scuff can be less than two percent of the total defect pixel area. A model that has not been specifically trained to respect that boundary will classify them interchangeably, which means your scratch trend data is contaminated with scuff events and vice versa.

For a QA Vision Lead, this is not an academic distinction. Scratch root causes — roll damage, strip guide misalignment, debris in the roll bite — are completely different from scuff root causes — excessive inter-stand tension, roll surface finish degradation, improper coil handling. If the classification is unreliable, the root cause analysis that depends on it is unreliable, and the process engineers stop trusting the system's output within weeks. The defect taxonomy must be designed from the outset around class boundaries that map to distinct root causes and distinct corrective actions, not just distinct visual appearances.

The six-class taxonomy below represents the defect categories that iFactory has validated across multiple cold-rolling installations, with per-class accuracy rates measured against human-expert-labeled datasets of more than 50,000 defect instances per class. Each class is defined by its physical characteristics, its root cause category, and the specific imaging conditions required for reliable detection.

Full Defect Class Map — Cold-Rolled Sheet

Each class below includes the baseline accuracy rate achieved on the iFactory reference dataset, the typical false-positive rate observed across deployed installations, and the root cause category that makes the class actionable for process engineers. The accuracy bars represent classification precision — when the model says a defect is class X, how often it is actually class X.

Edge Dent
MECHANICAL
Precision

94.2%
False Positive Rate: 1.8%

Localized deformation at the strip edge caused by coil handling, improper uncoiler tension, or impact from guide rollers. Typically appears as a smooth, rounded indentation 5 to 30 millimeters in length along the edge 10 to 50 millimeters from the strip boundary. Detected using angled edge lighting that creates shadow contrast in the deformation zone. Root cause category: mechanical handling and tension control.

Oil Stain
CHEMICAL
Precision

91.3%
False Positive Rate: 3.4%

Residual rolling oil or emulsion that is not uniformly distributed across the strip surface, appearing as irregularly shaped dark patches with diffuse boundaries. Severity ranges from light film variations that evaporate during downstream annealing to heavy pooling that causes carbon deposition and surface discoloration. Detected using diffuse axial lighting at specific wavelengths where oil absorbs differently than bare metal. Root cause category: lubrication system and emulsion spray configuration.

Scratch
MECHANICAL
Precision

93.1%
False Positive Rate: 2.1%

Linear groove cut into the strip surface by a foreign object, a damaged work roll, or a misaligned component in the strip path. Characterized by sharp edges, consistent width along the length, and depth that penetrates below the surface oxide layer. Length ranges from a few millimeters to the full strip length for continuous roll-contact scratches. Detected using directional lighting aligned perpendicular to the expected scratch orientation. Root cause category: roll condition, strip path alignment, and debris exclusion.

Scuff
MECHANICAL
Precision

86.7%
False Positive Rate: 5.6%

Superficial abrasion that displaces surface material without cutting into the substrate, typically caused by excessive friction between the strip and roll surfaces or between the strip and guide components. Visually similar to scratches but with diffuse edges, variable width, and no consistent depth profile. This is the hardest class in the taxonomy due to its visual overlap with scratches and its sensitivity to lighting angle. Root cause category: inter-stand tension, roll surface finish degradation, and guide contact pressure.

Herringbone
STRUCTURAL
Precision

97.8%
False Positive Rate: 0.4%

Characteristic cross-hatched pattern of surface ridges oriented at approximately 45 degrees to the rolling direction, caused by work roll chatter vibration or by strip slip between stands during speed transitions. The pattern is highly regular and visually distinctive, which makes it the easiest class for AI to identify with high precision. Severity is graded by ridge height and the percentage of strip width affected. Root cause category: mill vibration, chatter control, and inter-stand speed synchronization.

Edge Cracking
STRUCTURAL
Precision

94.0%
False Positive Rate: 2.2%

Fractures initiating at the strip edge and propagating inward, caused by excessive reduction per pass, insufficient edge trimming, or material with marginal edge ductility. Cracks range from hairline fractures visible only under magnification to open splits several centimeters long. Detected using edge-specific camera stations with backlighting that creates high contrast at the crack boundary. Root cause category: reduction schedule, edge trim configuration, and incoming material edge quality.

The Confusion Matrix — Where Classification Breaks Down

A headline accuracy number like 93 percent tells you nothing about where the model fails. The confusion matrix below shows exactly how each defect class is misclassified — which classes get confused with which other classes, and at what rate. This is the document a QA Vision Lead needs to evaluate whether a model is ready for production use, because it reveals the specific confusion pairs that will contaminate your defect trend data if left unaddressed.

Edge Dent Oil Stain Scratch Scuff Herringbone Edge Crack
Edge Dent 94.2 0.8 1.2 0.5 0.3 3.0
Oil Stain 0.5 91.3 0.3 2.8 0.4 4.7
Scratch 0.4 0.2 93.1 5.2 0.3 0.8
Scuff 0.6 1.1 8.3 86.7 0.9 2.4
Herringbone 0.2 0.1 0.4 0.6 97.8 0.9
Edge Crack 2.1 0.9 0.7 1.8 0.5 94.0

Correct Classification (Diagonal)

0.0 - 1.0% Misclassification

1.1 - 3.0% Misclassification

3.1 - 5.0% Misclassification

5.1 - 8.0% Misclassification

8.1%+ Misclassification

What Drives the Three Critical Confusion Pairs

Three confusion pairs in the matrix above account for the vast majority of classification errors, and understanding the physical basis for each pair is essential for deciding whether to accept the current performance or invest in plant-specific retraining. Each pair below explains why the model confuses the two classes and what plant-specific factors can reduce the confusion rate.

Scratch misclassified as Scuff: 5.2%
Scratch to Scuff

This is the single largest confusion pair in the taxonomy. At line speed, shallow scratches — where the groove depth is less than 5 microns — produce surface reflections that are nearly indistinguishable from scuff marks under standard lighting. The model relies on edge sharpness and width consistency along the defect length to distinguish the two classes, but when the scratch is shallow and the strip surface has minor roughness variation, those features become unreliable. Plant-specific retraining with shallow scratch samples from your specific roll grades and strip alloys typically reduces this confusion from 5.2 percent to below 3 percent by teaching the model the subtle depth cues that are unique to your surface finish range.

Scuff misclassified as Scratch: 8.3%
Scuff to Scratch

The reverse confusion is even larger because scuff marks are inherently more variable in appearance than scratches. A scuff that occurs at a guide roller where the contact pressure is uneven can produce a linear abrasion pattern that mimics a scratch in both shape and orientation. The model classifies these as scratches because the linear morphology matches the scratch template more closely than the diffuse scuff template. This confusion is particularly prevalent on lines where guide roller surfaces are worn or where strip tension fluctuates during speed changes, because both conditions produce scuff patterns that look more linear than the training data expects. Adding scuff samples from your specific guide roller configurations to the training set is the most effective correction.

Oil Stain misclassified as Edge Crack: 4.7%
Oil Stain to Edge Crack

This confusion occurs when oil pooling concentrates near the strip edge and creates a dark, irregular boundary that the model interprets as a crack opening. The spectral signature of pooled oil and the shadow pattern at a crack edge share enough low-frequency features to cause misclassification when the oil stain is large and the edge camera's lighting angle creates partial shadow effects at the oil boundary. Plants that run heavy emulsion applications or that have marginal edge wiping systems see this confusion pair more frequently. The correction is a combination of adding edge-region oil stain samples to the training set and adjusting the edge camera lighting angle to reduce shadow overlap with the oil boundary.

Per-Class Detection Complexity and False-Positive Drivers

False positives are more damaging to system adoption than false negatives in most cold-rolling environments, because a false positive triggers unnecessary downgrading or reinspection that costs real money, while a false negative on a minor defect may be caught by downstream inspection. The cards below break down what specifically drives false positives for each class, so QA teams can assess whether their plant's operating conditions are likely to produce higher or lower FP rates than the baseline.

Edge Dent
FP Driver: Edge waviness mimicking dent shadows

Edge waviness caused by strip shape control issues creates shadow patterns under angled lighting that resemble dent profiles. The model can distinguish a true dent from waviness when the shadow gradient is steep enough, but mild waviness with gradual depth transitions falls within the dent feature space. Plants with ongoing shape control issues will see FP rates above the 1.8 percent baseline. Adjusting the edge lighting angle to reduce waviness shadow intensity without losing dent contrast is the primary mitigation.

Oil Stain
FP Driver: Emulsion pattern variation and water spots

Normal emulsion drainage patterns on the strip surface create variations in reflectivity that can trigger the oil stain class when the pattern is locally concentrated. Water spots from imperfect drying also produce dark patches that overlap with the oil stain feature space. The FP rate is highly sensitive to the emulsion concentration and the drying system performance. Plants running high emulsion ratios or with aging dryer nozzles typically see FP rates of 4 to 6 percent on this class until the model is retrained with samples that include normal emulsion variation patterns specific to that plant.

Scratch
FP Driver: Roll marks and chatter lines

Periodic roll marks and light chatter lines produce linear features on the strip surface that share morphological characteristics with scratches. The model distinguishes scratches from roll marks primarily by periodicity — roll marks repeat at the roll circumference interval, while scratches are typically singular or randomly spaced. When the roll mark periodicity is irregular due to roll eccentricity or when a single roll mark is isolated within the camera field of view, the periodicity cue is lost and the model classifies it as a scratch. Including plant-specific roll mark samples with known periodicity characteristics in the training set resolves most of these false positives.

Scuff
FP Driver: Surface roughness variation and coating irregularities

Cold-rolled strip surface roughness varies across the width due to roll crown profiles and backup roll flattening patterns. Areas of naturally higher roughness scatter light differently and can produce a mottled appearance that triggers the scuff class. Temporary coating applied for corrosion protection during storage also creates surface texture variations that overlap with scuff features. The scuff class has the highest baseline FP rate at 5.6 percent precisely because it is defined by the absence of sharp features rather than the presence of distinctive ones. Retraining with roughness map data from your specific roll campaigns is the most effective reduction strategy.

Herringbone
FP Driver: Virtually none — most distinctive pattern class

The cross-hatched geometry of herringbone defects is so visually distinctive that false positives on this class are rare at 0.4 percent. The few false positives that do occur are typically caused by directional surface finishing marks from backup rolls that happen to approximate the herringbone angle, or by light interference patterns from overlapping lighting sources. This class is the safest starting point for validating a new installation because any detections are almost certainly real, which builds operator confidence in the system before moving to the harder classes.

Edge Cracking
FP Driver: Edge trim burr and slitter marks

Edge trimming and slitting operations leave micro-burr and shear marks along the strip edge that create discontinuities in the edge profile detected by the edge camera. When the burr height is above 10 microns and the lighting angle creates a shadow on the burr side, the edge profile shows a discontinuity that the model interprets as a crack initiation point. Plants with dull slitter knives or excessive trimmer knife clearance see elevated FP rates on this class. Retraining with edge profiles that include burr patterns at various knife wear stages, combined with a knife change schedule that keeps burr height below the detection threshold, brings the FP rate back to baseline.

Plant-Specific Retraining — Why Generic Models Fail on Your Line

The baseline accuracy and confusion matrix numbers published above are measured on iFactory's reference dataset, which includes defect samples from multiple cold-rolling installations across different steel grades, roll configurations, and lighting setups. That dataset is deliberately broad to ensure the model generalizes well across the industry. But no generic dataset captures the specific combination of surface finish characteristics, lighting geometry, defect morphology variations, and false-positive trigger conditions that exist on your specific line. Plants that deploy the generic model without retraining typically see overall accuracy 4 to 7 percentage points below the published baseline, with the gap concentrated in the hardest classes — scuff and oil stain — where plant-specific variation matters most.

01
Collect Plant-Specific Samples
02
Expert Labeling and Review
03
Augmented Retraining Cycle
04
Validated Deployment

The retraining process requires a minimum of 2,000 labeled samples per class from your specific line to produce a statistically meaningful improvement. Those samples must include not only the defect instances but also the false-positive trigger conditions — edge waviness, emulsion patterns, roll marks, burr profiles — that are specific to your plant. The expert labeling step is where most plants underestimate the effort: every sample must be reviewed and classified by a senior QA engineer who understands the physical distinction between classes, not just the visual one. A sample that looks like a scuff to a junior inspector may be correctly identified as a shallow scratch by an engineer who knows the roll condition history at the time the sample was collected. That domain expertise is what turns a retrained model from slightly better to significantly better.

See the Full Confusion Matrix for Your Strip Grades and Line Speed
iFactory publishes per-class accuracy and confusion data before you commit — including the specific confusion pairs that matter most for your defect mix and surface finish range.

Before and After Plant-Specific Retraining

The table below shows representative accuracy improvement from plant-specific retraining at a 5-stand tandem cold mill processing 0.4 to 2.0 mm strip in low-carbon and IF steel grades. The before column reflects the generic model performance in the first two weeks of shadow-mode operation. The after column reflects performance after retraining with 3,200 plant-specific samples per class and four weeks of additional validation. The improvement is not uniform — it concentrates in exactly the classes where plant-specific variation drives false positives.

Defect Class Generic Model Precision Retrained Precision Improvement FP Rate Before FP Rate After
Edge Dent 89.4% 95.1% +5.7 pts 4.2% 1.4%
Oil Stain 84.7% 93.6% +8.9 pts 7.1% 2.8%
Scratch 88.9% 94.8% +5.9 pts 4.8% 1.6%
Scuff 79.3% 89.4% +10.1 pts 9.8% 4.2%
Herringbone 96.1% 98.2% +2.1 pts 0.9% 0.3%
Edge Cracking 90.2% 95.3% +5.1 pts 3.9% 1.7%
Why False-Positive Rate Matters More Than Headline Accuracy for QA Adoption

A model with 94 percent overall accuracy and a 6 percent false-positive rate will be disabled by operators within three months. A model with 89 percent overall accuracy and a 1.5 percent false-positive rate will be adopted and gradually improved. The reason is purely operational: every false positive requires someone to walk to the line, inspect the flagged location, determine that the defect is not real, and clear the alarm. At high line speeds with frequent false positives, this activity consumes more operator time than the manual inspection it was supposed to replace. QA Vision Leads who have lived through failed AI deployments consistently identify false-positive fatigue as the number one reason operators stop trusting the system. The published per-class false-positive rates in iFactory's defect taxonomy exist specifically so that you can evaluate the FP exposure before deployment, set realistic expectations with your operations team, and prioritize the retraining effort on the classes where FP reduction will have the largest operational impact.

Frequently Asked Questions

Can the defect taxonomy be extended to include additional classes beyond the six published?
Yes. The six-class taxonomy covers the most common and highest-impact defect categories for cold-rolled sheet, but individual mills may have defect types specific to their product mix, processing route, or customer requirements — for example, coil break, pinhole, inclusion, or specific coating defects for galvanized products. iFactory's platform supports taxonomy extension by adding new class definitions, collecting and labeling samples for the new classes, and retraining the model with the expanded class set. The critical requirement is that the new class must be visually and physically distinct from existing classes — adding a class that overlaps significantly with an existing class will degrade accuracy on both. Book a demo to discuss which additional classes make sense for your product range.
How many samples are needed to retrain the model for a new steel grade?
The minimum viable sample count for retraining on a new grade is 1,500 to 2,000 labeled instances per defect class, collected from production runs that span the full range of thickness, width, and surface finish variations for that grade. For grades that are visually similar to existing trained grades — for example, moving from DC01 to DC03 low-carbon steel — the improvement from retraining may be marginal and the existing model may perform adequately. For grades with significantly different surface characteristics — for example, moving from low-carbon to ferritic stainless or to high-strength low-alloy grades — the retraining benefit is substantial because the surface reflectivity, roughness profile, and defect appearance all change enough to push the generic model outside its trained feature space. Contact support for grade-specific retraining guidance.
Does the confusion matrix get updated after plant-specific retraining?
Yes. After each retraining cycle and validation period, iFactory generates an updated confusion matrix specific to your installation that reflects the model's actual performance on your defect distribution, your strip grades, and your lighting conditions. This plant-specific matrix replaces the generic baseline as the performance reference for your QA team. The updated matrix also serves as the input for deciding whether additional retraining iterations are warranted — if a specific confusion pair remains above your acceptance threshold after retraining, the matrix tells you exactly which class pair needs more targeted samples and whether the issue is a training data gap or a fundamental visual ambiguity that requires a different approach such as additional sensor modalities or modified lighting geometry. Schedule a demo to see how the matrix update workflow operates in practice.
How does the system handle defects that fall between two classes?
Every classification output includes not only the predicted class but also a confidence score and a probability distribution across all six classes. When a defect falls in the ambiguous zone between two classes — for example, a defect that the model scores as 48 percent scratch and 42 percent scuff — the system flags it as a low-confidence classification rather than forcing it into a single class. Low-confidence classifications are routed to the QA review queue rather than triggering automatic disposition actions, which prevents ambiguous defects from contaminating the trend data for either class. Over time, the accumulation of low-confidence samples provides a prioritized list of edge cases that should be added to the next retraining cycle to sharpen the boundary between the confused classes. Reach out to support to discuss low-confidence handling configurations for your workflow.
What is the typical timeline from generic deployment to retrained production use?
The full cycle from generic model deployment to retrained production use typically runs 10 to 16 weeks. Weeks one through four are shadow-mode operation with the generic model, during which the system collects defect images and classification outputs without affecting production decisions. Weeks three through six overlap with expert labeling of the collected samples — this starts before shadow mode ends because the labeling pipeline takes time to build momentum. Weeks six through ten are the actual retraining and internal validation, where the retrained model is tested against a held-out set of plant-specific samples that were not used in training. Weeks ten through twelve are a second shadow-mode period with the retrained model, during which the updated confusion matrix is generated and compared against the acceptance criteria. Weeks twelve through sixteen cover any additional retraining iterations needed to bring the remaining confusion pairs below threshold, followed by the final go-live decision. Book a demo to get a detailed timeline customized to your mill's sampling capacity and defect frequency.
QA VISION LEADS · CONFUSION MATRIX PUBLISHED · RETRAINABLE PER PLANT
Get the Per-Class Accuracy Data Before You Commit to a Vision Platform
See the full six-class defect taxonomy, confusion matrix, and false-positive rates for cold-rolled sheet — and what plant-specific retraining does to each number.

Share This Story, Choose Your Platform!