Your Model Scored 92%. How Stable Is That Number?

Published by Industry AI Decision

Articles / AI Learning / Machine Learning & Data

Your model scored 92% on a test set. Is that number stable? Not necessarily. A single score is a point estimate from one sample. Pair it with an interval that shows how much the score could move if the represented cases were sampled again.

One approachable way to estimate that movement is the bootstrap: repeatedly rebuild test samples from the cases you already observed, calculate the score each time, and examine the resulting spread. It does not make a weak test set representative, but it makes uncertainty in that test set visible.

Think of the test set as a jar of result cards

Imagine that every held-out inspection case is a card marked “correct” or “wrong.” Put all the cards in a jar. Draw the same number of cards, returning each card before the next draw. Because a card can appear more than once—or not at all—the score changes slightly.

That is resampling with replacement. Repeat it many times and the collection of scores shows the wobble that one 92% result hides. The ModernDive open textbook explains this distinction clearly: a sampling distribution would use many samples from the population, while a bootstrap distribution repeatedly resamples the one sample available.

A worked inspection example

Consider a fictional vision model tested on 200 held-out part images. It classifies 184 correctly, so the observed accuracy is 92%. We resample those 200 recorded outcomes with replacement 1,000 times and recalculate accuracy after every resample.

Histogram of 1,000 bootstrap accuracy scores from a synthetic 200-case test set, centered near 92 percent with a middle 95 percent interval from 88.0 to 95.5 percent
Figure 1. Data / Chart. Synthetic teaching data: the middle 95% of 1,000 resampled scores runs from 88.0% to 95.5%. Fixed random seed: 20261005.

In this fixed-seed teaching run, the middle 95% of resampled scores is 88.0% to 95.5%. That range is the interval. It does not mean the model has a 95% chance of scoring inside that range on every future production batch. It says that the observed test set supports more uncertainty than the point score alone suggests.

The sampling unit matters

The simple row-by-row bootstrap assumes the 200 outcomes act like independent cases. Factory data often violate that assumption. Ten images of one part, frames from the same video, or units from the same batch may share lighting, material and process conditions.

If outcomes share a batch, resample whole batches rather than individual rows. Otherwise the interval can look narrower than the evidence deserves. Also keep the test set untouched by model selection: repeated tuning against the same test results turns them into training feedback and weakens the final claim.

Four-step decision framework: report point score and interval, check the sampling unit, match resampling to rows or batches, then decide whether to collect more evidence or test distribution shift
Figure 2. Decision Framework. Report the interval, respect batch dependence, then choose the next evidence step.

What a wider interval should change

A wide interval is not a failed model. It is a signal that the current evaluation cannot support a precise claim. The next decision may be to label more independent cases, add missing production conditions, or redesign the split around batches, lots or time periods.

A narrow interval answers only a smaller question: the score is stable within the represented test conditions. It does not test line changes, sensor drift, a new supplier, label errors or rare defects that never appeared in the sample.

Three takeaways

  1. A single test score hides sampling uncertainty; report a point score and an interval together.
  2. Resample the unit that can vary independently. For correlated factory data, that may be a batch rather than an image.
  3. A narrow interval does not prove future production performance or protection from distribution shift.

One action

Before the next model review, ask the analyst to run 1,000 bootstrap resamples on the untouched test set, using batches when outcomes are grouped. Put the point score, interval, sample size and resampling unit on the same slide.

Limitation

The basic percentile bootstrap used here is a teaching method, not a universal guarantee. Tiny or biased samples, rare outcomes, strong dependence and some statistics can require other interval methods or a better data-collection design.

Answer to the opening question: 92% is stable only if its interval is acceptably narrow for the decision and the test cases represent the production conditions you care about.

Sources

Next lesson: Should You Trust One Test Split? Cross-Validation, Explained

PUT THE IDEAS TO WORK

Assess a workflow from your own operation.

Choose a calculator or review for business value, OEE, capacity, equipment, integration or AI governance. Save your assumptions and results in a private workspace.

KEEP READING

Recommended next reads.

Continue with three pieces chosen for the topic you are reading.

ARTICLE EMAILS

Get the next article by email.

Subscribe for new Industry AI Decision analysis and learning articles. Email subscription is separate from a free member account.

You can unsubscribe or change delivery preferences at any time. Privacy policy

Leave a Reply

Discover more from Industry AI Decision | Agentic Manufacturing & Decision Intelligence

Subscribe now to keep reading and get access to the full archive.

Continue reading