Can an AI quality model learn from inspection records if operators label only suspicious parts? Not safely—not until you check what the uninspected output looks like. The records may describe the inspection rule more clearly than the production line.
This problem is called selection bias: the data you observe differ systematically from the data you do not observe. Google’s Machine Learning Glossary identifies incomplete coverage as one form of selection bias. In manufacturing, that can happen when a vision system or an operator decides which parts deserve a manual check.
Articles / AI Learning / AI Fundamentals
The label rule changes what the model sees
Imagine a line making 1,000 parts during one shift. A camera flags 80 as suspicious, so operators inspect those 80 and record 20 defects and 60 acceptable parts. The other 920 continue without a manual label.

The 80 reviewed parts form the labeled sample. It is useful, but it was not drawn evenly from the shift. It was selected because the first system already considered those parts unusual. Treating the 920 uninspected parts as “good” would replace missing evidence with an assumption.
A limited analogy helps: judging an entire warehouse from the boxes already placed in the exception cage tells you a lot about exceptions, not necessarily about every box.
What the 25% result really means
In this fictional example, 20 of the 80 inspected parts are defective, so the observed defect rate is 25% among flagged parts. The arithmetic is correct. The scope is narrow.
Dividing 20 by all 1,000 parts to report a 2% line defect rate would not be justified. The outcomes of 920 parts remain unknown. Some may be good; some may be defects the first inspection rule missed. The missing labels are not random.

This matters during both learning and evaluation. A model trained only on selected records may learn the characteristics of already-flagged parts. A test built the same way can look accurate while never measuring ordinary output. During inference, the model can still produce a score for every part, but a score does not reveal the missing real-world outcomes.
Add an audit sample
The practical repair is an audit sample: manually inspect a small random set from the parts that were not flagged. NIST’s guidance on random sampling explains why representative samples support valid conclusions, while nonrepresentative samples can lead to the wrong model. NIST also advises industrial AI teams to check that available data cover the full intended operating scope and real-world variation.
For ten shifts, keep inspecting every flagged part, but also randomly select 50 unflagged parts per shift. Record the two groups separately. Compare their defect rates, materials, machines, products and operating conditions. Do not combine the groups into one headline rate until the sampling design and any weighting are reviewed.
This audit does more than estimate missed defects. It shows where the current labeled data have thin coverage. That evidence can guide the next labeling budget and expose whether one product family, shift or machine is almost absent from the training set.
A feedback loop can deepen the gap
Dive into Deep Learning warns that deployed systems can influence the data later collected about them. Here, the model flags parts, people label mostly those parts, and the next model learns from that selected history. Without an independent audit stream, the system can keep confirming its own attention pattern.
Three takeaways
- A correct percentage can still answer the wrong population question.
- Uninspected parts are unknown, not proven acceptable.
- A small random audit can reveal coverage gaps that exception-only labels hide.
One action
For the next ten shifts, randomly inspect 50 unflagged parts per shift and compare them with the flagged group. Keep the two sampling routes visible in the dataset.
After those outcomes exist, the site’s threshold trade-off tool can compare missed detections and false alarms. It is a follow-up check, not a substitute for collecting the missing labels.
Limitation
A random audit reduces one source of selection bias; it does not automatically make the sample representative of every product, season, supplier or failure mode. Rare safety-critical defects may require a different, risk-based sampling plan.
Answer: inspection records can support learning, but only within the population they actually represent. If labels come mostly from suspicious parts, add independent evidence from unflagged output before making line-wide claims.
Sources
- Aston Zhang, Zachary C. Lipton, Mu Li and Alexander J. Smola, Dive into Deep Learning, 1st ed. (Cambridge University Press, 2023), Chapter 4 §4.7, “Environment and Distribution Shift”; accessed September 29, 2026.
- Google for Developers, Machine Learning Glossary: Selection bias; accessed September 29, 2026.
- NIST, “NIST Researcher Describes Data Considerations for Industrial Artificial Intelligence”, published February 1, 2025 and updated March 17, 2025; accessed September 29, 2026.
- NIST/SEMATECH e-Handbook, “Sample Types and Their Applications”; accessed September 29, 2026.
Next lesson: Should You Flip Every Inspection Image?
Leave a Reply