Motion & Gain
Commercialization AI Transformation Work Insights Expertise Subscribe Contact Subscribe →
Home / Insights / Challenge & Resolution
Challenge & Resolution

The Override Gap

Consumer goods companies bought the demand-planning models. The returns showed up only where they redesigned the decision around them.

Motion&Gain 28 July 2026 5 min read
The gap is not a modelling error. It is a decision that arrived one cycle late.

A demand plan is not a forecast. It is a chain of decisions — what to make, what to hold, what to promise a retailer on Thursday — and the forecast is only the first link. Over the last five years, every large consumer goods company has upgraded that first link. Far fewer have touched the rest of the chain. The difference between the two groups is most of the value.

The Challenge

By the early 2020s, the standard planning stack in consumer goods had stopped holding. A statistical baseline — exponential smoothing, ARIMA, and their descendants — was generated on a monthly cycle, then adjusted by hand before it reached supply.

It broke in two places, and only one of them was the model.

The model was blind to why demand moved. History-based methods rest on a single assumption: that the best predictor of future demand is past demand. That holds for a stable SKU in a stable channel. It fails precisely where the money is. A promotion, a heatwave, a competitor's price cut, a retailer's planogram reset — by the time any of these appear in the sales record, the event is over. The forecast is reporting yesterday's news at the exact moment it needs to be forward-looking. For manufacturers where a third of volume moves on promotion, that is not an edge case. That is the business.

The override was unmeasured. Planners knew the model was wrong, so they corrected it. Nobody measured whether the corrections helped. When one manufacturer ran forecast value added retrospectively across fourteen months of manual adjustments, roughly 40% of them moved weighted error in the wrong direction. The organization was paying skilled people to degrade its own forecast — and had no instrument that would have told it.

That is the override gap: the distance between what the system could have known and what the organization actually decided.

Why buying a better model didn't close it?

The obvious fix — a more sophisticated algorithm — has a poor track record on its own. Three reasons recur.

Opacity breeds override. A model that produces a number without a traceable reason invites correction. The research found more than half of supply chain planning teams cite explainability as a meaningful barrier to adoption. Unexplained forecasts get overwritten, and every overwrite resets the trust clock.

Sophistication is not accuracy. Forecast error is frequently a data problem wearing a modelling costume. Fragmented master data, unflagged product transitions, missing promotional calendars — no ensemble compensates for inputs that never described the business correctly.

The planning calendar was never re-cut. A model that can re-forecast daily, plugged into a process that decides monthly, delivers a monthly decision. The cadence, not the algorithm, sets the ceiling.

The Resolution

The companies that got paid made four decisions, not one. They are unglamorous and they are sequential.

1. Stop running one model across the whole portfolio. Segment the SKU base by forecastability, then match the method. The 2026 landscape has settled into three usable layers: gradient boosting with engineered features (LightGBM, XGBoost) on high-velocity items where covariates are rich; time-series foundation models — Amazon's Chronos-2, Google's TimesFM, Salesforce's Moirai, Nixtla's TimeGPT — used zero-shot or lightly fine-tuned across the long tail and new product introductions where there is no history to learn from; and classical methods retained where series are genuinely stable. The honest read on foundation models is that zero-shot performance sits between classical methods and fully trained deep models — impressive for zero training, not yet dominant. Fine-tuning on a few hundred domain points closes most of the gap. Treat them as a way to cover the tail cheaply, not as a replacement for the core.

2. Feed the model causes, not just history. This is where the accuracy actually comes from. Price and promotional mechanic, competitor activity, media weight, weather, retailer POS sell-out, local events, macro indicators. Unilever's weather-driven models lifted forecast accuracy for ice cream in Sweden by around 10 percentage points — a category where a six-day temperature swing is the demand driver and the sales history is noise. Nestlé's Brazilian data team reached roughly 94% accuracy predicting sell-out over twelve-week horizons by modelling drivers rather than extrapolating.

3. Shorten the loop. Move from a monthly consensus cycle to continuous re-forecasting, with retraining triggered by events rather than the calendar. The caution here is real: over-retraining on noisy data fits transient variance and increases error. Weekly retraining earns its keep when external signals move faster than a month; monthly is sufficient for stable, high-volume SKUs. Anchor the schedule to a value-added check before new weights go to production.

4. Govern the human — don't remove them. This is the move most often skipped, and it is the one that converts model accuracy into business outcome. Make forecast value added a standing governance metric, not a diagnostic: measure the naïve baseline, the machine forecast, the planner override and the consensus number separately, by SKU and channel. Publish who adds value and where. Then automate what the data says is safe — Gartner's guidance for 2026 puts touchless forecasting for stable SKUs and automated replenishment parameters squarely in the sweet spot — and route the remainder to a planner with the evidence attached. The planner stops re-typing numbers and starts adjudicating exceptions.

What lies ahead...

Agentic planning, for now, in most of the portfolio. Gartner's 2026 assessment is blunt: vendors claiming end-to-end autonomous supply chain planning before 2027 are overstating what is possible, and "agent washing" is a live procurement risk. Cross-enterprise negotiation and dynamic cost trade-offs are poor early candidates. High-volume, low-cost-of-error decisions are good ones. Start there, keep the audit trail, keep the hand-off points explicit.

Any programme that starts with the tool. The sequencing that works is data foundation, then segmentation, then model, then cadence, then governance. Reverse it and the pilot produces a better number that nobody uses.

The Bottom Line

Every large consumer goods company now has access to the same models. Ninety-one percent of supply chain leaders say they will use AI for demand forecasting within two years, which means the model is no longer the differentiator — it is the entry fee.

What still differentiates is narrower and harder: whether the organization has measured its own override behavior, whether the planning cadence matches the signal speed, and whether inventory and service policy were re-cut to spend the accuracy that was bought.

The forecast was never the hard part. Closing the gap between what the model knows and what the business decides — that is the work.

Share
LinkedIn Email
Keep reading

Get challenges & resolutions by email

Real commercial problems and the specific decisions that resolved them. Free, and you choose what you receive.

Subscribe →