Articles / AI Learning / Machine Learning & Data
Your model scored 92% on a test set. Is that number stable? Not necessarily. A single score is a point estimate from one sample. Pair it with an interval that shows how much the score could move if the represented cases were sampled again.
One approachable way to estimate that movement is the bootstrap: repeatedly rebuild test samples from the cases you already observed, calculate the score each time, and examine the resulting spread. It does not make a weak test set representative, but it makes uncertainty in that test set visible.
Think of the test set as a jar of result cards
Imagine that every held-out inspection case is a card marked “correct” or “wrong.” Put all the cards in a jar. Draw the same number of cards, returning each card before the next draw. Because a card can appear more than once—or not at all—the score changes slightly.
That is resampling with replacement. Repeat it many times and the collection of scores shows the wobble that one 92% result hides. The ModernDive open textbook explains this distinction clearly: a sampling distribution would use many samples from the population, while a bootstrap distribution repeatedly resamples the one sample available.
A worked inspection example
Consider a fictional vision model tested on 200 held-out part images. It classifies 184 correctly, so the observed accuracy is 92%. We resample those 200 recorded outcomes with replacement 1,000 times and recalculate accuracy after every resample.

In this fixed-seed teaching run, the middle 95% of resampled scores is 88.0% to 95.5%. That range is the interval. It does not mean the model has a 95% chance of scoring inside that range on every future production batch. It says that the observed test set supports more uncertainty than the point score alone suggests.
The sampling unit matters
The simple row-by-row bootstrap assumes the 200 outcomes act like independent cases. Factory data often violate that assumption. Ten images of one part, frames from the same video, or units from the same batch may share lighting, material and process conditions.
If outcomes share a batch, resample whole batches rather than individual rows. Otherwise the interval can look narrower than the evidence deserves. Also keep the test set untouched by model selection: repeated tuning against the same test results turns them into training feedback and weakens the final claim.

What a wider interval should change
A wide interval is not a failed model. It is a signal that the current evaluation cannot support a precise claim. The next decision may be to label more independent cases, add missing production conditions, or redesign the split around batches, lots or time periods.
A narrow interval answers only a smaller question: the score is stable within the represented test conditions. It does not test line changes, sensor drift, a new supplier, label errors or rare defects that never appeared in the sample.
Three takeaways
- A single test score hides sampling uncertainty; report a point score and an interval together.
- Resample the unit that can vary independently. For correlated factory data, that may be a batch rather than an image.
- A narrow interval does not prove future production performance or protection from distribution shift.
One action
Before the next model review, ask the analyst to run 1,000 bootstrap resamples on the untouched test set, using batches when outcomes are grouped. Put the point score, interval, sample size and resampling unit on the same slide.
Limitation
The basic percentile bootstrap used here is a teaching method, not a universal guarantee. Tiny or biased samples, rare outcomes, strong dependence and some statistics can require other interval methods or a better data-collection design.
Answer to the opening question: 92% is stable only if its interval is acceptably narrow for the decision and the test cases represent the production conditions you care about.
Sources
- Chester Ismay, Albert Y. Kim and Arturo Valdivia, Statistical Inference via Data Science: A ModernDive into R and the Tidyverse, Second Edition (2026), Chapter 8, especially Sections 8.2–8.3. Main conceptual source; accessed 5 October 2026.
- Bradley Efron, “Bootstrap Methods: Another Look at the Jackknife,” The Annals of Statistics 7(1), 1979. Original method paper; accessed 5 October 2026.
- NIST/SEMATECH e-Handbook, “Bootstrap Plot”. Resampling procedure and cautions; accessed 5 October 2026.
- NIST/SEMATECH e-Handbook, “What are confidence intervals?”. Interpretation of interval coverage; accessed 5 October 2026.
Next lesson: Should You Trust One Test Split? Cross-Validation, Explained
Leave a Reply