Inspection Data Quality: Are Only Suspicious Parts Labeled?

Published by Industry AI Decision

Can an AI quality model learn from inspection records if operators label only suspicious parts? Not safely—not until you check what the uninspected output looks like. The records may describe the inspection rule more clearly than the production line.

This problem is called selection bias: the data you observe differ systematically from the data you do not observe. Google’s Machine Learning Glossary identifies incomplete coverage as one form of selection bias. In manufacturing, that can happen when a vision system or an operator decides which parts deserve a manual check.

Articles / AI Learning / AI Fundamentals

The label rule changes what the model sees

Imagine a line making 1,000 parts during one shift. A camera flags 80 as suspicious, so operators inspect those 80 and record 20 defects and 60 acceptable parts. The other 920 continue without a manual label.

Annotated factory inspection scene showing 1,000 parts, 80 flagged for review and 920 unflagged parts with unknown labels
Fictional teaching example: the inspection rule determines which parts receive labels.

The 80 reviewed parts form the labeled sample. It is useful, but it was not drawn evenly from the shift. It was selected because the first system already considered those parts unusual. Treating the 920 uninspected parts as “good” would replace missing evidence with an assumption.

A limited analogy helps: judging an entire warehouse from the boxes already placed in the exception cage tells you a lot about exceptions, not necessarily about every box.

What the 25% result really means

In this fictional example, 20 of the 80 inspected parts are defective, so the observed defect rate is 25% among flagged parts. The arithmetic is correct. The scope is narrow.

Dividing 20 by all 1,000 parts to report a 2% line defect rate would not be justified. The outcomes of 920 parts remain unknown. Some may be good; some may be defects the first inspection rule missed. The missing labels are not random.

Comparison chart showing 20 defects among 80 inspected parts and 920 uninspected parts whose labels remain unknown
Fictional counts: 20 of 80 applies to the inspected sample, not automatically to all 1,000 parts.

This matters during both learning and evaluation. A model trained only on selected records may learn the characteristics of already-flagged parts. A test built the same way can look accurate while never measuring ordinary output. During inference, the model can still produce a score for every part, but a score does not reveal the missing real-world outcomes.

Add an audit sample

The practical repair is an audit sample: manually inspect a small random set from the parts that were not flagged. NIST’s guidance on random sampling explains why representative samples support valid conclusions, while nonrepresentative samples can lead to the wrong model. NIST also advises industrial AI teams to check that available data cover the full intended operating scope and real-world variation.

For ten shifts, keep inspecting every flagged part, but also randomly select 50 unflagged parts per shift. Record the two groups separately. Compare their defect rates, materials, machines, products and operating conditions. Do not combine the groups into one headline rate until the sampling design and any weighting are reviewed.

This audit does more than estimate missed defects. It shows where the current labeled data have thin coverage. That evidence can guide the next labeling budget and expose whether one product family, shift or machine is almost absent from the training set.

A feedback loop can deepen the gap

Dive into Deep Learning warns that deployed systems can influence the data later collected about them. Here, the model flags parts, people label mostly those parts, and the next model learns from that selected history. Without an independent audit stream, the system can keep confirming its own attention pattern.

Three takeaways

  • A correct percentage can still answer the wrong population question.
  • Uninspected parts are unknown, not proven acceptable.
  • A small random audit can reveal coverage gaps that exception-only labels hide.

One action

For the next ten shifts, randomly inspect 50 unflagged parts per shift and compare them with the flagged group. Keep the two sampling routes visible in the dataset.

After those outcomes exist, the site’s threshold trade-off tool can compare missed detections and false alarms. It is a follow-up check, not a substitute for collecting the missing labels.

Limitation

A random audit reduces one source of selection bias; it does not automatically make the sample representative of every product, season, supplier or failure mode. Rare safety-critical defects may require a different, risk-based sampling plan.

Answer: inspection records can support learning, but only within the population they actually represent. If labels come mostly from suspicious parts, add independent evidence from unflagged output before making line-wide claims.

Sources

Next lesson: Should You Flip Every Inspection Image?

PUT THE IDEAS TO WORK

Assess a workflow from your own operation.

Choose a calculator or review for business value, OEE, capacity, equipment, integration or AI governance. Save your assumptions and results in a private workspace.

KEEP READING

Recommended next reads.

Continue with three pieces chosen for the topic you are reading.

ARTICLE EMAILS

Get the next article by email.

Subscribe for new Industry AI Decision analysis and learning articles. Email subscription is separate from a free member account.

You can unsubscribe or change delivery preferences at any time. Privacy policy

Leave a Reply

Discover more from Industry AI Decision | Agentic Manufacturing & Decision Intelligence

Subscribe now to keep reading and get access to the full archive.

Continue reading