A factory AI model should abstain when the expected cost of an automatic mistake is greater than the cost of review—and only after its confidence has been tested on representative data. A score of 0.82 is not, by itself, permission to act.
Abstention, also called the reject option, adds a third outcome to classification: decide automatically, reject automatically, or say “not sure” and route the case to a defined reviewer. The point is not to make the model look cautious. It is to spend human attention where uncertainty and consequence are highest.
The useful trade-off is risk versus coverage
Christopher Bishop’s decision-theory treatment separates the model’s inference from the decision that follows. If the largest class probability does not clear a chosen threshold, the system can reject the automatic decision. Raising that threshold usually reduces errors among the cases the model handles, but it also sends more work to people.
Two measures make that trade-off visible:
- Coverage = automatically decided cases ÷ all eligible cases.
- Selective risk = errors among automatically decided cases ÷ automatically decided cases.
- Review rate = 1 − coverage.

A worked inspection example
Consider a fictional visual-inspection model applied to 1,000 machined parts. With no abstention, it decides all 1,000 cases and makes 60 mistakes: a 6% error rate.
At one candidate threshold, it decides 700 cases, makes 21 mistakes among them, and sends 300 cases to review. Coverage is 70%; selective risk is 21 ÷ 700 = 3%; review rate is 30%.
That is a real improvement only if the review path can absorb 300 cases without creating delay or superficial approvals. If the shift can review only 200, the policy is operationally infeasible even though the model’s automatic error rate looks better. The threshold therefore belongs to an operating decision, not only a model-tuning notebook.
Confidence must be calibrated before it becomes a gate
Many classifiers emit a probability-like number that is too confident or too conservative. A well-calibrated model that assigns roughly 0.8 to many comparable cases should be correct on about 80% of those cases. Calibration curves compare predicted probability with observed frequency; they do not guarantee that one individual case is correct.
This is why the abstention threshold should be evaluated on data that resembles the intended line, product, camera, shift, and defect mix. Reusing training data, or labeling only suspicious parts, can make the threshold appear safer than it is. See the lessons on probability calibration and selection bias in inspection data.
“Not sure” needs an owner and a reason code
A low maximum probability is only one reason to abstain. Missing sensor values, poor image quality, an unseen product variant, disagreement between models, or a case outside the approved operating range may also trigger review.

The review record should identify the reason, reviewer, final decision, response time, and whether the reviewed case later becomes useful labeled evidence. NIST’s AI Risk Management Framework emphasizes defined human–AI roles, documented knowledge limits, and measured oversight. It does not prescribe one universal confidence threshold.
What to measure before enabling abstention
- Plot risk and review workload across several candidate thresholds on untouched, deployment-like data.
- Check coverage and errors separately for important products, lines, shifts, and rare defect classes.
- Choose a threshold that fits both error tolerance and verified review capacity.
- Freeze it for a defined pilot window, then monitor drift, queue time, overrides, and missed defects.
Research on selective classification demonstrates the general risk–coverage idea on benchmark datasets. That evidence does not prove the same threshold or error reduction will hold in a factory. Human reviewers can also be inconsistent, and distribution shift can invalidate a previously acceptable policy.
Three takeaways
- Abstention is a third decision path, not a claim that confidence equals truth.
- Lower selective risk usually costs automation coverage and creates review work.
- The threshold must be validated with calibrated probabilities, deployment-like data, named ownership, and measurable capacity.
Sources
- Christopher M. Bishop, Pattern Recognition and Machine Learning (Springer, 2006), Chapter 1 §1.5.3, “The reject option.” Official Microsoft Research PDF.
- Yonatan Geifman and Ran El-Yaniv, “Selective Classification for Deep Neural Networks,” submitted 23 May 2017. arXiv.
- scikit-learn 1.9.1 documentation, “Probability calibration,” accessed 7 October 2026. Documentation.
- NIST, AI Risk Management Framework Core, Govern 3.2, Map 2.2 and Map 3.5, accessed 7 October 2026. AI RMF Core.
Next action
For the next 200 comparable decisions, record confidence, automatic or reviewed route, final outcome, review time, and override reason. Compare at least three frozen thresholds before choosing one operating policy.