A Factory Model Scores 94%. Which Line Is Failing?

Published by Industry AI Decision

Which line is failing? Line B. In the synthetic example below, a factory inspection model is 96% accurate on 900 Line A cases but only 76% accurate on 100 Line B cases. The volume-weighted overall score is still 94%. If the approval meeting sees only that aggregate, the smaller line disappears inside the average.

The remedy is slice evaluation: report the same model metric for operationally meaningful subsets such as line, product variant, shift, camera, supplier lot or defect family. An overall score describes the dataset mix. A slice score helps locate where a decision may be unsafe.

The aggregate is a weighted vote

Overall accuracy counts every labelled case equally. That sounds neutral, but high-volume conditions receive more influence because they contribute more cases. Here, Line A supplies 90% of the evaluation set:

  • Line A: 864 correct out of 900, or 96%
  • Line B: 76 correct out of 100, or 76%
  • Overall: (864 + 76) ÷ 1,000 = 94%

Nothing is mathematically wrong with 94%. The error is managerial: treating a dataset-weighted summary as proof that every deployment condition is acceptable.

Bar chart showing 96% accuracy on Line A, 94% overall, and 76% on Line B in a synthetic 1,000-case example.
Figure 1. Synthetic teaching data: a 94% volume-weighted score hides 76% accuracy on the lower-volume Line B.

Choose slices that can change an operating decision

A slice is useful when it maps to a different condition, owner or action. “All records from last Tuesday” may be easy to query but useless for operations. “Night-shift images from Camera 3 after the lighting retrofit” can point to a specific equipment or data problem.

For manufacturing, begin with a short, pre-declared slice set: production line, product family, shift, inspection station and a small number of known risk conditions. Add intersections only when they answer a real question. Crossing every field with every other field creates tiny groups and a dashboard full of noise.

A worked review: from 94% to a bounded action

Suppose the model screens cosmetic defects before final inspection. The 76% Line B result should trigger investigation, but it does not yet prove the model caused scrap or escapes. The team should first confirm that both lines used the same label definition, threshold and sampling window. Then compare camera settings, product mix and error types.

Sample size matters too. One hundred labelled cases can justify a closer look, but a precise pass/fail claim still needs uncertainty and a threshold tied to operational harm. Ten examples with one error should not be displayed as “90.0%” without a warning. If evidence is thin, report insufficient evidence, keep human review, and collect the next labelled sample.

Editorial diagram moving from an overall model score to line, product, shift, and sample slices, followed by evidence checks.
Figure 2. A practical review path: open deployment-relevant slices, then check sample size, uncertainty, harm threshold and ownership.

Build the review into the model release gate

  1. Declare the slices before reading results. Use the conditions the model will actually face, not whichever grouping produces the most dramatic chart.
  2. Show counts beside every metric. Pair accuracy, recall or false-negative rate with the number of labelled cases and, where practical, an uncertainty interval. The same discipline underlies bootstrap model-score stability checks.
  3. Assign a response owner. A weak slice should lead to one of four bounded actions: collect more evidence, repair data or equipment, adjust the operating threshold, or restrict deployment.

Slice analysis complements, rather than replaces, a sound test design. A slice cannot repair leakage, selection bias or a holdout that does not represent production. Review why one test split can mislead and how selection bias enters inspection data before treating any subgroup result as a deployment guarantee.

Three takeaways

  1. An overall score reflects the evaluation mix. High-volume conditions can dominate it.
  2. Slice metrics are decision aids, not automatic verdicts. Counts, uncertainty and harm thresholds still matter.
  3. The best slice set is operationally small. Choose dimensions with a clear owner and response.

Limitations

The 900/100 case example is synthetic and explains aggregation; it is not measured factory evidence. Accuracy is also not always the right metric. For rare critical defects, recall, false-negative rate and review workload may be more decision-relevant. Slice boundaries can create false confidence if labels are inconsistent, samples are selected after seeing outcomes, or many tiny slices are tested without accounting for chance variation.

Next action

Add one page to your next model review: overall metric, five pre-declared deployment slices, labelled-case count, uncertainty note, acceptance threshold and named owner. If a critical slice lacks enough evidence, keep the decision open rather than letting the aggregate close it.

Sources

Related reading

PUT THE IDEAS TO WORK

Assess a workflow from your own operation.

Choose a calculator or review for business value, OEE, capacity, equipment, integration or AI governance. Save your assumptions and results in a private workspace.

KEEP READING

Recommended next reads.

Continue with three pieces chosen for the topic you are reading.

ARTICLE EMAILS

Get the next article by email.

Subscribe for new Industry AI Decision analysis and learning articles. Email subscription is separate from a free member account.

You can unsubscribe or change delivery preferences at any time. Privacy policy

Leave a Reply

Discover more from Industry AI Decision | Agentic Manufacturing & Decision Intelligence

Subscribe now to keep reading and get access to the full archive.

Continue reading