Which line is failing? Line B. In the synthetic example below, a factory inspection model is 96% accurate on 900 Line A cases but only 76% accurate on 100 Line B cases. The volume-weighted overall score is still 94%. If the approval meeting sees only that aggregate, the smaller line disappears inside the average.
The remedy is slice evaluation: report the same model metric for operationally meaningful subsets such as line, product variant, shift, camera, supplier lot or defect family. An overall score describes the dataset mix. A slice score helps locate where a decision may be unsafe.
The aggregate is a weighted vote
Overall accuracy counts every labelled case equally. That sounds neutral, but high-volume conditions receive more influence because they contribute more cases. Here, Line A supplies 90% of the evaluation set:
- Line A: 864 correct out of 900, or 96%
- Line B: 76 correct out of 100, or 76%
- Overall: (864 + 76) ÷ 1,000 = 94%
Nothing is mathematically wrong with 94%. The error is managerial: treating a dataset-weighted summary as proof that every deployment condition is acceptable.

Choose slices that can change an operating decision
A slice is useful when it maps to a different condition, owner or action. “All records from last Tuesday” may be easy to query but useless for operations. “Night-shift images from Camera 3 after the lighting retrofit” can point to a specific equipment or data problem.
For manufacturing, begin with a short, pre-declared slice set: production line, product family, shift, inspection station and a small number of known risk conditions. Add intersections only when they answer a real question. Crossing every field with every other field creates tiny groups and a dashboard full of noise.
A worked review: from 94% to a bounded action
Suppose the model screens cosmetic defects before final inspection. The 76% Line B result should trigger investigation, but it does not yet prove the model caused scrap or escapes. The team should first confirm that both lines used the same label definition, threshold and sampling window. Then compare camera settings, product mix and error types.
Sample size matters too. One hundred labelled cases can justify a closer look, but a precise pass/fail claim still needs uncertainty and a threshold tied to operational harm. Ten examples with one error should not be displayed as “90.0%” without a warning. If evidence is thin, report insufficient evidence, keep human review, and collect the next labelled sample.

Build the review into the model release gate
- Declare the slices before reading results. Use the conditions the model will actually face, not whichever grouping produces the most dramatic chart.
- Show counts beside every metric. Pair accuracy, recall or false-negative rate with the number of labelled cases and, where practical, an uncertainty interval. The same discipline underlies bootstrap model-score stability checks.
- Assign a response owner. A weak slice should lead to one of four bounded actions: collect more evidence, repair data or equipment, adjust the operating threshold, or restrict deployment.
Slice analysis complements, rather than replaces, a sound test design. A slice cannot repair leakage, selection bias or a holdout that does not represent production. Review why one test split can mislead and how selection bias enters inspection data before treating any subgroup result as a deployment guarantee.
Three takeaways
- An overall score reflects the evaluation mix. High-volume conditions can dominate it.
- Slice metrics are decision aids, not automatic verdicts. Counts, uncertainty and harm thresholds still matter.
- The best slice set is operationally small. Choose dimensions with a clear owner and response.
Limitations
The 900/100 case example is synthetic and explains aggregation; it is not measured factory evidence. Accuracy is also not always the right metric. For rare critical defects, recall, false-negative rate and review workload may be more decision-relevant. Slice boundaries can create false confidence if labels are inconsistent, samples are selected after seeing outcomes, or many tiny slices are tested without accounting for chance variation.
Next action
Add one page to your next model review: overall metric, five pre-declared deployment slices, labelled-case count, uncertainty note, acceptance threshold and named owner. If a critical slice lacks enough evidence, keep the decision open rather than letting the aggregate close it.
Sources
- Vijay Janapa Reddi, Introduction to Machine Learning Systems, Volume I: Foundations, v0.7.2 (2026), ML Operations, “Slice analysis”. The open textbook explains how aggregate production metrics can mask severe degradation in a subpopulation.
- TensorFlow, “TensorFlow Model Analysis Setup”, last updated January 28, 2021. Official documentation for configuring overall, feature-key and crossed-feature evaluation slices.
- NIST, AI Risk Management Framework 1.0, Core Measure 2.3 and 2.5, 2023. Performance should be demonstrated for deployment-like conditions, and generalizability limits documented.
- Vincent S. Chen et al., “Slice-based Learning: A Programming Model for Residual Learning in Critical Data Slices”, NeurIPS 2019; revised February 29, 2020. The paper defines critical subsets as slices and documents how coarse metrics can hide poor subset performance.
Leave a Reply