AI Demand Forecasting, Explained: How the Models Work, What Accuracy to Expect, and What It Takes to Run
AI demand forecasting explained: what the models learn from, gradient boosting vs foundation time-series models, the accuracy to expect, the data you need, and what it costs.
Written by Max Zeshut
Founder at Agentmelt
TL;DR: AI demand forecasting is a model trained on your sales history plus the signals a spreadsheet cannot hold — promotions, price, weather, holidays, marketing, web traffic — producing a forecast per product, per location, per week, with an error band. On weekly SKU-location data it typically cuts forecast error by 20–40% against a moving average or last year's number, and more where promotions drive demand. It is worth running above roughly 200 SKU-locations; below that, a good baseline with promotion adjustments does most of the work. The forecast is only half the system: the other half is the weekly loop that scores last week's forecast against actuals and explains what moved, so planners review exceptions instead of rebuilding the sheet.
Buildable version: the demand forecasting workflow — what arrives, what happens, who approves, a free template, and the price to have it run for you.
What "AI demand forecasting" actually means
Three things are sold under the phrase, and they are not the same:
- A statistical forecast with a better name. Exponential smoothing, ARIMA, seasonal decomposition — the methods in every planning system for thirty years. Reliable on stable, seasonal products; blind to anything that is not in the past sales curve.
- A machine-learning forecast. A model trained on your history and on external drivers, learning how each driver moves demand for each product. This is what most vendors mean and what this article is about.
- A large language model reading the forecast. Not forecasting at all, but useful: turning the numbers into "sales of the 500 ml variant will fall 18% in week 41 because the promotion ends and the competitor's launches" — the paragraph a planner would write, produced for 800 products every Monday.
A working system uses 2 for the numbers and 3 for the words. Anyone selling 3 as the forecast is selling a chatbot.
What the model learns from
The gain over a spreadsheet comes almost entirely from what the model is allowed to see. In rough order of how much each usually moves the error:
- Sales history per product and location — at least two years for seasonality, weekly grain for most B2B and retail.
- Promotions and price: the promotion calendar, discount depth, price changes, the competitor's price where you have it. For consumer goods this is the single largest driver and the one spreadsheets handle worst.
- Calendar effects: holidays, paydays, school terms, events, the number of trading days in the period.
- Product attributes: category, brand, pack size, launch date — so a new product borrows the curve of its siblings until it has history of its own.
- External signals: weather (for anything seasonal or perishable), web traffic and search interest, macro indicators for capital goods, marketing spend.
- Supply-side facts: stockouts. A product that sold zero because it was unavailable did not have zero demand, and a model trained on that row learns the wrong lesson. Censoring stockout weeks is the most common fix that vendors skip.
How the models work
Gradient-boosted trees (LightGBM, XGBoost, CatBoost) are the workhorse. Each product-week becomes a row of features — the lags of its own sales, the calendar, the promotion flags, the attributes — and the model learns the mapping from features to next week's demand across all products at once. That is the key difference from classical methods, which fit one curve per product: a boosted model learns that "20% off in the second week of a month" lifts demand by a certain shape, from every product that ever had it, and applies it to a product that never has. They train in minutes on a laptop for tens of thousands of series.
Foundation time-series models (Chronos, TimesFM, TimeGPT and their successors) are pre-trained on millions of public series and forecast yours with no training at all, the way a language model completes text. They are strong on short histories and new products, weaker at using your promotion and price drivers unless the vendor supports covariates. In 2026 the practical pattern is to run a foundation model as a second opinion and for cold starts, with a boosted model as the primary once you have history.
Deep models (N-BEATS, Temporal Fusion Transformers, DeepAR) sit between: trained on your data, able to learn complex interactions, more expensive to run and to explain. Worth it at large scale with rich drivers; rarely the first choice.
Whichever model produces the point forecast, two additions turn it into something a planner can use:
- Quantiles. A forecast of 1,200 units with a 10th–90th percentile band of 900–1,600 says how far to trust it, and the band is what safety stock should be computed from — not a percentage everyone agreed on.
- Hierarchical reconciliation. Forecasts by SKU, by category and by region should add up. Reconciliation forces them to, and usually improves every level in the process.
Accuracy: what to expect, and how to measure it honestly
The honest measure is the one the workflow reports every week, not a benchmark on a slide. Two cautions before any number:
- MAPE lies on intermittent demand. A product that sells 0, 0, 3, 0 has infinite percentage errors. Use WAPE (weighted absolute percentage error — total absolute error over total actual sales) or a scaled error, and report it by segment: fast movers, slow movers, new products, promoted weeks.
- Compare with the baseline you actually had. The gain is against your spreadsheet or your ERP's moving average, measured on the same weeks. Vendors quote gains against a naive forecast nobody uses.
With those caveats, the typical result on weekly SKU-location data for a company with decent history and a promotion calendar:
| Segment | Baseline WAPE (moving average / last year) | ML forecast WAPE | What moves it |
|---|---|---|---|
| Fast-moving, stable products | 20–30% | 12–20% | Calendar and price effects |
| Promoted products | 40–60% | 20–35% | The promotion calendar, discount depth |
| Slow / intermittent | 60–100% | 45–80% | Quantile forecasts; better to plan service level than the point |
| New products (first 8 weeks) | no baseline | 35–60% | Attribute-based cold start, foundation models |
A 20–40% relative reduction in error at the SKU-location level is the common range. What it is worth depends on what the error costs you: stockouts on the fast movers, write-offs on the perishables, working capital everywhere else.
The weekly loop is the product
A forecast run once is a project; a forecast run every week with a scorecard is a system. The loop in the demand forecasting workflow:
- Pull last week's sales, stock positions and the updated promotion calendar from the ERP or the planning tool.
- Score last week's forecast against actuals by segment; write the accuracy into the record.
- Retrain or refresh the model (weekly retrain for boosted models is cheap; foundation models need none).
- Forecast the next 8–13 weeks per SKU-location with quantiles; reconcile the hierarchy.
- Explain: a language model reads the forecast, the drivers and the changes since last week and writes the exception list — which products moved, why, and what needs a planner's eye.
- Deliver into the planning system or the sheet planners already use, with the overrides recorded separately so their accuracy can be scored too.
Step 5 is where the time saving lives. A planner with 800 products cannot read 800 forecasts; they can read twelve exceptions with reasons.
What it takes to run
Data: two years of weekly sales per SKU-location (one year works, with a weaker seasonal picture), the promotion and price history, product attributes, and stockout flags if you have them. If the ERP cannot give you weekly history, a nightly export to a table is the first step, and it is a day's work, not a project.
Infrastructure: for boosted models, a scheduled job on a small server or a hosted model endpoint; foundation models are called as an API. Neither needs a data science team to run once the pipeline is built; someone needs to own the weekly scorecard.
People: a planner who reviews the exceptions and keeps the right to overrule, and one person who owns the promotion calendar's accuracy — the model is only as good as the calendar it is told about.
Cost: as an installed workflow on this site, $297 a month for one ERP source and up to 5,000 SKU-locations, hosted model included; custom models, planning-system integrations or multiple entities are a custom build. A planning platform's forecasting module starts at tens of thousands a year and makes sense when you are also buying the rest of the platform.
When not to bother
- Fewer than about 200 SKU-locations: a good baseline with promotion adjustments in a spreadsheet gets most of the value.
- No promotion calendar and no external drivers: the ML model has little to learn beyond seasonality, and a statistical method does that already.
- Demand you control, not predict: made-to-order or contract volumes belong in a schedule, not a forecast.
- A forecast nobody acts on: if reorder points, production plans or budgets do not read the forecast, accuracy is a vanity metric. Start with inventory optimisation tied to the forecast, or do not start.
Questions, answered
How does AI demand forecasting work?
A model is trained on your sales history together with the drivers that move demand — promotions, price, holidays, weather, marketing, product attributes — and learns, across all products at once, how each driver changes next week's sales. It then forecasts each product at each location for the coming weeks with an error band. Run weekly, it scores its own past accuracy and a language model explains which forecasts moved and why, so a planner reviews exceptions rather than every number.
How accurate is AI demand forecasting?
On weekly SKU-location data, 20–40% less error than a moving average or last year's figure, measured as WAPE by segment; more where promotions dominate, less on intermittent demand where the honest output is a probability band rather than a point. The number that matters is the one your own weekly scorecard reports against the baseline you used before.
What is the difference between AI, ML and statistical demand forecasting?
Statistical methods fit one curve per product from its own history. Machine learning trains one model across all products on history plus external drivers, and is what "AI demand forecasting" usually means. A large language model does not forecast; it reads the forecast and writes the explanation. Working systems combine the second and the third.
Which AI demand forecasting tools should we look at?
Open-source: LightGBM or XGBoost with a feature pipeline, Nixtla's libraries for classical and neural methods, and foundation models such as Chronos and TimesFM for cold starts. Hosted: TimeGPT and the forecasting modules of planning platforms. The tool matters less than the pipeline around it — the weekly scorecard, the promotion calendar and the delivery into the system planners use — which is what the workflow here packages.
Can AI demand forecasting handle new products?
Partly. A new product borrows the demand curve of products with the same attributes until it has eight to twelve weeks of history, and foundation models help because they have seen thousands of launch curves. Expect roughly double the error of an established product in the first two months, and plan service levels from the quantile band rather than the point.