A tattooed person pointing at finance charts and graphs on a whiteboard. Forecasting mistakes seen from the repair side
Photo by https://kaboompics.com/ on Pexels

Maintenance

Part of Grading a forecasting cycle honestly before anyone trusts it

Forecasting mistakes seen from the repair side

Forecasting failures that survive a good model: single numbers, targets in disguise, missing baselines, unlogged overrides and forecasts nobody scores.

Forecasts fail in a small number of recognizable ways, and almost none of them are failures of technique. They are failures of framing, of incentive, or of record keeping.

Each entry below names the move, what it costs, and the smaller thing to do instead. All nine survive a perfectly reasonable model.

What to take away

  • The most damaging errors happen before and after the model runs, not inside it.
  • A forecast that cannot be scored later is not a forecast. It is an opinion with decimal places.
  • Where a forecast and a target are the same number, you no longer have a forecast, and no amount of method will recover one.

1. Publishing a single number

A point estimate with no range reads as a commitment. Everyone downstream plans against it as though it were certain, and the first miss becomes a credibility problem rather than an expected outcome.

The fix is to publish an interval and to say in words what the interval means. If a stakeholder insists on one number, give them the point and the range in the same sentence rather than in a footnote.

2. Letting the forecast become the target

Once the forecast is what someone is held to, it stops describing the future and starts describing what is safe to say. Sales forecasts drift low, capacity forecasts drift high, and the direction tells you exactly where the pressure is.

This is Goodhart's law operating on a prediction rather than on a metric: the measure has become a goal, so it stops measuring.

The fix is to keep the forecast and the commitment as separate numbers, both visible, with the gap between them named. The gap is management information in its own right.

3. No baseline, so no way to tell whether the model helps

A model produces a number, the number is used, and nobody ever computes what a repeat of last period would have given.

The fix is to run the naive baseline every cycle and store it beside the forecast. It costs almost nothing and it settles the question of whether the modeling effort earns its keep.

4. Grading the model on the data it was fitted to

In-sample performance always looks good, and improving it always looks like progress. What is actually being improved is the fit to noise that will not repeat.

Overfitting is easy to describe and hard to resist, because the feedback loop rewards it every time.

The fix is a holdout reserved before fitting, plus evaluation rolled across several origins rather than one.

5. Re-specifying the model after every miss

A period goes badly, so a term is added to explain it. Next quarter something else goes wrong, so another term appears. After a year the model explains history perfectly and forecasts nothing.

This is the Texas sharpshooter fallacy run as a process: the target is drawn after the shots, every time.

The fix is a change rule: model changes happen on a schedule, are justified against holdout performance, and are logged. Explaining last month is not a justification.

6. Overrides with no record

Judgment adjustments are often right and always undocumented. Six months later nobody can say whether the model was wrong, the override was wrong, or both.

The fix is to store the pre-override number and score all three: the model, the override, and the naive baseline. Teams are frequently surprised by which one wins.

7. Forecasting further than the decision needs

A twelve-month forecast for a decision made six weeks out produces ten months of numbers that exist only to be argued about, and error grows the whole way.

The fix is to set the horizon from the lead time of the decision and to stop there, publishing longer views only as a separately labeled scenario.

8. Using a driver you also have to forecast

Regressing on a variable that is itself unknown at the time of the forecast quietly doubles your uncertainty and hides it inside a model that reports only its own error.

The fix is to restrict drivers to things known in advance: calendar effects, contracted volumes, published schedules. Anything else becomes a scenario input, stated as an assumption.

9. Never scoring anything

The forecast is produced, the actuals arrive, and nobody puts them side by side. The cycle repeats for years, and the organization has no idea whether it forecasts well.

The fix is a standing review with a fixed date, a fixed measure, and a written record of what changed as a result. Twenty minutes a cycle is enough.

The pattern underneath

Seven of the nine are failures to keep a record: no baseline stored, no pre-override number, no holdout, no change log, no score. Forecasting is unusual among analytical work in that the truth arrives on its own schedule, which means the entire opportunity to learn depends on having written down what you believed before you found out.

That is a filing problem, not a statistical one, and it can be fixed by a team with no modeling skill at all.

Related reading on this site

The method and the review discipline are in how forecasts are built and reviewed. For the metric design habits that keep the second mistake visible, see metric shapes and their traps. For the sponsor-level version of the same incentive problem, see nine business intelligence mistakes.

The record keeping the closing section argues for is a governance question, covered in data governance. The publication discipline around a published number sits in building a report worth keeping.

Common questions

Our forecast has been the plan for years. How do we separate them?

Publish both for two cycles without proposing any change, and let the gap be visible. The conversation about which one to hold people to happens by itself once the two numbers are on the same page.

Are judgment overrides bad?

No. They are frequently the most valuable input, particularly around events the model cannot see. What is bad is applying them without keeping the number they replaced.

How do we score a forecast for something that only happens a few times a year?

Score the direction and the interval coverage rather than the level, and accept that it will take several cycles to say anything. State that limitation when you publish.

We changed the model and accuracy improved. Is that enough?

Only if the improvement holds on data the new model never saw, and only if the baseline did not improve by the same amount for its own reasons. Both checks are quick and both are usually skipped.

More in Maintenance

Latest from Practice Desk

Costs

IRS tax checklist for US analytics contractors, from 1099s to quarterly taxes

Business analytics and intelligence contractors file Schedule C, pay quarterly estimated taxes, handle 1099-NEC forms, and claim deductions. Here is the checklist.

Industry

Chicago logistics and retail metrics, a dashboard overview for the Midwest

Dashboard metrics for Chicago logistics and retail teams cover shipment, inventory and consumer data pulled from TMS, WMS, POS and BLS sources.

Industry

How Boston biotech and healthcare teams build intelligence dashboards

An intelligence dashboard in Boston biotech and healthcare pairs clinical, trial and claims metrics with strict data governance rules and city dashboard patterns.

Industry

Research Triangle Park analytics hiring and the university pipeline behind it

Intelligence dashboard metrics hiring in Research Triangle Park runs on life sciences and tech employers, university degree programs, and posted metrics roles.