Articles / AI Learning / Practical AI Skills
If a test image shows a physical part that already appeared in training, the score may not measure performance on a new part. For most factory-inspection pilots, split by physical unit, wafer, lot or production run before augmentation or resampling.
The file can be new while the underlying case is not. One bracket may produce a top view, side view, close-up and reinspection image. A random image split can place some views in training and another view in test. The model then receives evidence about the same unit on both sides of the boundary.
Different files can still be the same case
This failure is called group leakage. The group is the real-world entity that should remain intact: one part, patient, machine, wafer, lot, household or production run. The correct group depends on the deployment question.
Vijay Janapa Reddi’s open textbook Introduction to Machine Learning Systems treats dataset partitions as trust boundaries. It explicitly lists duplicates, augmented variants of one source record and records from the same group appearing across splits as leakage. Google’s Machine Learning Crash Course likewise warns that duplicates in train and test can make evaluation look better than performance on unseen data.

Choose the boundary from deployment
Ask what must genuinely be unseen when the model is used. If the system will inspect newly manufactured parts, keep every image of a physical part in one partition. If the business question concerns a new supplier lot, group by lot. If the next month’s production is the target, preserve time order as explained in the point-in-time test for industrial AI.
A stricter boundary can be justified when variation is nested. Images may belong to parts, parts to lots and lots to one tool or shift. Splitting only by part prevents the same part from crossing the boundary, but it may still let one lot’s distinctive surface or one camera’s fixed pattern appear everywhere. State which boundary the score supports instead of calling every holdout “unseen.”
Worked example: 1,000 images, only 200 parts
This is a synthetic teaching example, not factory performance data. A team photographs 200 physical parts from five angles, producing 1,000 image files. It assigns 140 parts to training, 30 to validation and 30 to final test. Because every part has five images, the partitions contain 700, 150 and 150 images.
The arithmetic is simple; the evidence rule matters more. All five images from one part must travel together. The team checks that the intersection of part IDs between every pair of partitions is empty. Only then does the test represent 30 unseen parts rather than 150 conveniently new files.
If the team uses cross-validation, it should use a group-aware method. Scikit-learn’s GroupKFold keeps each group out of the training set when that group is used for testing. The broader lesson on why one test split is not enough still applies, but ordinary random folds cannot repair the wrong grouping rule.
Four gates before training

- Name the source entity. Add a stable unit, wafer, lot, tool or run identifier to the dataset manifest.
- Assign groups, not files. Put the complete group into train, validation or test.
- Augment after splitting. Derived crops, rotations and brightness variants stay inside their source partition. The guide to valid inspection-image transformations explains the separate question of whether an augmentation preserves the label.
- Freeze and audit. Check exact hashes, near-duplicates, group-ID intersections and timestamps before the final evaluation.
What group splitting does not fix
A zero-overlap group check does not prove deployment readiness. Preprocessing fitted on the full dataset can still leak information. Labels may include future outcomes. Repeated model selection can tune decisions to the final test. And a clean holdout cannot represent a product, defect or operating condition that was never collected.
The chosen group can also be wrong for the actual decision. If deployment intentionally compares later images of a known unit with its own earlier images, holding out entire units tests a stricter question. Document the target population and the unit of independence so reviewers know what the score does—and does not—support.
Three takeaways
- A new image file is not necessarily an independent test case.
- Split by the real entity that must be unseen at deployment, then create derived samples.
- Audit group IDs, duplicates and time boundaries before interpreting the model score.
Sources
- Vijay Janapa Reddi, Introduction to Machine Learning Systems, Volume I: Foundations, Data Engineering — Dataset Compilation, version 0.7.2, updated August 31, 2026.
- Google for Developers, “Datasets: Dividing the original dataset”, Machine Learning Crash Course, updated December 3, 2025.
- scikit-learn 1.9.1, GroupKFold documentation, accessed October 11, 2026.
- Sayash Kapoor and Arvind Narayanan, “Leakage and the Reproducibility Crisis in ML-based Science”, submitted July 14, 2022. Cross-domain evidence on leakage risks; not manufacturing-specific performance evidence.
Next action
Take the manifest for one inspection dataset and add three columns: source_entity_id, split and derived_from_id. Compute the intersection of source IDs across train, validation and test. Any non-zero intersection is a defect to resolve before the next score review.
Leave a Reply