
Reviews
Part of Grading a forecasting cycle honestly before anyone trusts it
Forecasting checklist, tested against experience as of 2027
Forecasting checks to run before publishing: a baseline, a holdout, calendar effects, intervals defined in words, logged overrides and an expiry rule.
A forecast leaves the building carrying more authority than it deserves. Once a number is in a plan, its assumptions are invisible and its uncertainty is gone.
These twelve checks are the ones worth running in the hour before you publish. They are ordered by how expensive the mistake is to fix afterwards.
What to take away
- Half of these are about what you say around the number, not about the model. That is where forecasts actually fail.
- A forecast with no stated baseline and no stated interval cannot be evaluated later, which means it cannot be improved.
- Write the checks into the publication process. Anything left to memory gets skipped in the week you most need it.
Before the model runs
1. Name the decision and the horizon it needs.
A forecast for a decision made in six weeks does not need twelve monthly points. Extending the horizon past the decision adds error and invites arguments about periods nobody will act on.
2. Fix the population and the unit before looking at history.
Which entities, which geography, which product set, and what counts as one observation. Changing this after seeing results is how a forecast becomes a preference.
3. Write down the naive baseline you have to beat.
Last period repeated, or the same period last year repeated. If the model cannot beat that, you have learned something real and cheap. The NIST handbook's introduction to time series analysis sets out the same starting discipline: characterize the series before fitting anything to it.
4. Reserve a holdout period and do not look at it.
Decide the split before fitting. A model evaluated on the data it was fitted to will always look good and will teach you nothing.
While the model runs
5. Plot the raw series and look at it.
Level shifts, gaps, a change in collection method, a period of zeros that means the feed broke rather than that demand stopped. Every one of these is visible in thirty seconds and invisible in a summary statistic.
6. Identify the calendar effects explicitly.
Weekday patterns, month length, holidays that move, a promotion window. Handle them or state that you have not, but do not let the model absorb them silently.
7. Check the error on the holdout, against the baseline.
Report both numbers. A model that beats the baseline by a small margin is often the right answer, and a model that loses to it should not be published.
8. Backtest across more than one window.
One holdout can flatter a model by accident. Rolling the evaluation across several origins is what backtesting is for, and it separates a model that works from a model that worked once.
Before it goes out
9. Publish an interval, and say what the interval means.
A prediction interval is a statement about where a future observation is likely to fall, and it is not the same thing as a confidence interval about a parameter. Readers conflate them constantly, so define the one you are using in the same sentence you present it.
10. Say what a probability attached to the forecast means.
If you publish a chance rather than a level, spell out what it is a chance of. Weather services have had to be unusually careful here, and the plain explanation of what a probability of precipitation means is a good model for the kind of sentence an internal forecast needs.
11. Record every judgment override, with a reason and an owner.
Overrides are legitimate. Unrecorded overrides make the forecast unevaluable, because next quarter nobody can tell whether the model or the person was wrong.
12. State the conditions under which this forecast stops being valid.
A named event, a threshold, or a date. Without it, a forecast made under one set of assumptions gets quoted for a year after those assumptions expired.
A short version to keep beside the process
| Stage | The check | The failure it prevents |
|---|---|---|
| Scope | Decision and horizon named | Forecasting periods nobody acts on |
| Scope | Population fixed in writing | Retrofitting the definition to the result |
| Fit | Baseline written down first | Complexity with nothing to justify it |
| Fit | Holdout reserved before fitting | A model graded on its own homework |
| Fit | Raw series inspected | Treating a broken feed as a real trend |
| Publish | Interval defined in words | A point estimate read as a promise |
| Publish | Overrides logged with owners | An unevaluable forecast next quarter |
| Publish | Expiry condition stated | A stale forecast quoted as current |
What this checklist deliberately leaves out
It says nothing about which method to use, because the method matters far less than the framing. A simple method with an honest interval and a recorded baseline is more useful than a sophisticated one published as a single number.
It also says nothing about accuracy targets. A target for forecast accuracy set without reference to the series volatility is a wish, and it usually produces smoothed forecasts rather than better ones.
Related reading on this site
The full method, including when not to forecast at all, is in how forecasts are built and reviewed. Publication habits carry over from building a report worth keeping, and the wording that should accompany a miss is in writing the paragraph beside the numbers.
The definitional work behind check two sits in writing a metric definition, and the analysis sequence behind check five is in how to work through a question.
Common questions
We forecast fifty series every month. Can this run on all of them?
Checks three, four, seven and eleven can be automated across a whole portfolio. The rest are per-series judgment, so run them on the handful of series that carry real money and let the others run on the automated subset.
Our stakeholders reject intervals and want one number.
Give them the number and put the interval beside it in the same line rather than in a footnote. Resistance usually falls away once people see that the interval is narrow for the series they care about most.
Is beating the naive baseline a low bar?
It is a low bar that a surprising share of published forecasts do not clear, particularly on series with strong seasonality where the seasonal naive comparison is the honest one.
How long should we keep the override log?
At least as long as the horizon you forecast, plus one review cycle. The log is only useful at the moment somebody asks why last year's number was wrong, and that question always arrives late.




