Articles / AI Learning / Machine Learning & Data
A point forecast alone cannot tell a planner how certain the estimate is. Ask for a prediction interval: a lower and upper bound designed to contain a stated share of comparable future outcomes. If a cycle-time model predicts 120 minutes with an 80% interval of 105–140 minutes, the useful answer is the range, not just 120.
The 80% label is not a guarantee for this one job. It is a repeated-forecast target: if the model and operating conditions remain appropriate, roughly 80% of comparable outcomes should fall inside their intervals over many forecasts.
Three terms are enough
- Point forecast: one expected value, such as 120 minutes.
- Prediction interval: a range for a future outcome, such as 105–140 minutes.
- Coverage: the share of actual outcomes that land inside their intervals.
Think of a delivery window. “12:00” is a point; “11:45–12:20” is a range. The window is not a promise, but it lets the receiving team decide whether one person or a full unloading crew must wait.

What changes in the factory decision?
Consider a fictional line that predicts the cycle time for a custom production job. The model reports 120 minutes. A downstream cell plans to start its setup at minute 125.
That schedule looks safe if the team sees only the point forecast. The 105–140 minute interval reveals a different risk: a meaningful set of similar jobs may still be running after minute 125. The planner can now compare the cost of a short idle period with the cost of an interrupted changeover. The model does not make that trade-off; it supplies evidence for it.
A wider interval is not automatically worse. It may honestly reflect a mixed product family, longer forecast horizon or unstable process. A narrow interval is not automatically better either. It can be overconfident and miss too many outcomes.
Audit the next 50 comparable jobs
Suppose the team records the lower bound, upper bound and actual cycle time for the next 50 comparable jobs. Thirty-nine actual times fall inside their intervals and 11 fall outside. The observed coverage is 39 ÷ 50 = 78%.
That is close to the 80% target, but 50 jobs are too few to declare the interval calibrated. The team should also record the average interval width. In this teaching example it is 35 minutes. An interval can hit its coverage target simply by becoming so wide that it no longer helps a scheduling decision.

How a model produces the range
During model development, the team can estimate a forecast distribution, learn lower and upper quantiles, or build intervals from past residuals. At inference time, the system returns the point and bounds for each new job. The scikit-learn example uses 5th- and 95th-percentile models to form a 90% interval; other methods make different assumptions.
The method matters less to an operating review than the test discipline. Use held-out or later data, keep the target and forecast horizon fixed, and segment results by product, shift or route when those conditions differ. A single random split can hide instability, so the existing lesson Can You Trust One Test Split? is a useful next check.
Three takeaways
- A point forecast answers “what is expected”; a prediction interval shows how uncertain that operating estimate is.
- Coverage and width must be reviewed together. Hitting 80% with an unusably wide range is not a practical win.
- Audit intervals on later, comparable cases and by operating segment. Training-set coverage is not deployment evidence.
One action
For the next 50 comparable forecasts, log the point, lower bound, upper bound, actual outcome and whether the actual lands inside. Report observed coverage and average width by product, shift and forecast horizon. Pair that review with MAE in real operating units so the team sees both typical point error and uncertainty.
Limitation
A prediction interval reflects the data and assumptions used to build it. It may not cover a new supplier, a machine breakdown, a policy change or another condition absent from the reference data. Recheck coverage after process changes; do not treat “80%” as an 80% guarantee for every individual job.
Answer to the opening question: the 120-minute forecast is only as useful as its tested range. Read 105–140 minutes as the decision input, then verify that the interval achieves useful coverage without becoming too wide.
Sources
- Rob J Hyndman and George Athanasopoulos, Forecasting: Principles and Practice, 3rd edition (OTexts, 2021), Chapter 5 §5.5, “Distributional forecasts and prediction intervals.” Online book updated 28 September 2026; accessed 4 October 2026.
- Hyndman and Athanasopoulos, Chapter 5 §5.9, “Evaluating distributional forecast accuracy,” especially interval width and penalties for missed outcomes. Accessed 4 October 2026.
- scikit-learn 1.9.1 documentation, “Prediction Intervals for Gradient Boosting Regression,” including quantile bounds and held-out coverage checks. Accessed 4 October 2026.
Next lesson: MAE Explained: How Far Are Manufacturing Predictions from Reality?
Leave a Reply