All documentation

Platform

Forecast Accuracy & Why You Can Trust the Data

Every forecast page in Market Lens shows you a number. The only question that matters about it is how much weight that number deserves — in your budget, in a hedge, in a position — and no amount of confident presentation can answer that for you. So the platform is built the other way round: you are never asked to trust a forecast blindly, because the historical record of that exact commodity, at that exact horizon, is published where you can go and read it first. Audit it, then decide. This page is the instruction manual for doing that.

Out-of-sample
Scored on prices it never saw
Five windows
Rolling backtests behind each score
Per commodity
Filterable by grain and horizon

The Forecast Scorecard: out-of-sample, not flattering

The Forecast Scorecard grades forecasts that were frozen at the moment they were made and then compared against the prices that actually happened afterwards. Nothing is scored on data it was fitted to, which is the distinction that decides whether a track record means anything: a model shown the answer can describe the past beautifully and still be useless about next quarter. Each score is built from five rolling backtest windows rather than one lucky stretch of market, so a calm quarter cannot carry the result — the discipline behind that is set out in How Market Lens Works: Data, Models & Backtesting.

TabThe question it answersWhat you read
Point AccuracyHow close was the predicted number?MAPE, RMSE and MAE. Lower is better on all three.
CalibrationDid the confidence band tell the truth?Coverage — how often the real price landed inside the band — against the level the band claims, plus the band's width.
RegimesDid it call the direction right?The directional hit-rate, and a Brier score for the quality of its probability calls. Lower is better on the Brier score.

A grain toggle switches between the monthly and daily records — grain being simply which of the two you are reading — and a horizon filter narrows to the exact lead time you rely on — which matters more than it sounds, because judging a 90-day call by its one-day accuracy tells you almost nothing about either. Every tab carries sortable per-commodity tables, so you can vet the single market you care about instead of accepting a platform-wide summary.

  1. Open the Scorecard for your commodity Not for the platform, and not for the category — for the specific commodity whose forecast is about to influence a decision. Predictability varies enormously between markets, and the average across them is nobody's actual situation.
  2. Filter to the horizon you actually use If you budget a year ahead, read the longest horizon. If you are timing a purchase this week, read the shortest. A model can be strong at one and ordinary at the other, and only one of those facts is relevant to you today.
  3. Check Point Accuracy This is the headline: how far the forecast landed from the price, on average, over the backtest windows. Lower is better. Read it as the typical size of the miss you should expect to live with, not as a promise about the next one.
  4. Check Calibration coverage against the band's stated promise The band claims a 90% range, so coverage close to that level means the range you plan against is honest. Coverage running well below it means the bands have been too optimistic for that commodity, and you should widen your own cushion accordingly.
  5. Check the direction record on the Regimes tab For decisions that hinge on which way a market moves rather than by how much — whether to hedge, whether to wait — the directional hit-rate and the Brier score are the columns that matter. A hit-rate meaningfully better than a coin flip is what tells you the signal is real.

What the measures mean in plain terms

  • MAPE — mean absolute percentage error: the average percentage the forecast was off, in either direction. It reads like a business KPI, which is why it is the headline figure. Lower is better.
  • MAE and RMSE — the same idea expressed in price units rather than percent. RMSE punishes large misses harder than MAE does, so an RMSE sitting well above the MAE for a commodity is a hint that this market occasionally throws big surprises even when it is usually well behaved. Lower is better on both.
  • Coverage — the calibration check. A 90% band should contain the real price about 90% of the time; coverage measures whether it did. This is a definition, not a claim: the measured value is what you go and look up.
  • Directional hit-rate — how often the model got the direction right. Better than a coin flip is the bar it has to clear before a directional call deserves any weight at all.
  • Brier score — a grade for probability calls, such as “70% chance of an up month”. It rewards being confident and correct, and penalises being confident and wrong. Lower is better.

Calibration is the quiet one worth understanding properly. If coverage matches the level the band claims, then the uncertainty range you have been planning against is honest — and the budget cushion you sized from that range is right-sized rather than wishful. A slightly larger average error with well-behaved bands is often more useful to a planner than a tighter central line whose range cannot be trusted, because it is the range you actually commit money against. And all of this varies by commodity and by horizon: some markets are inherently more predictable than others, and the Scorecard is where that difference stops being an abstraction.

Accuracy differs commodity by commodity and horizon by horizon, so there is no single number to quote and this page deliberately does not quote one. Check the Scorecard for the specific commodity and the specific horizon before you lean on its forecast; signed-in users will find it in the sidebar.

Where the numbers come from

A track record is only as meaningful as the prices it was scored against, so the inputs matter as much as the method. Every figure in Market Lens is built from independent, authoritative sources, and the classes below share a property that matters for auditing: each one is something you or an auditor can trace back to the body that issued it.

Type of dataWhy you can check it
Official statistics from public agenciesCompiled on a published methodology by bodies whose figures are public record, so a series can be reconciled against the institution that issued it.
Market and futures pricesTraded prices and the forward curve of contract prices — the same prices your counterparties see, which makes them the hardest input to get quietly wrong.
Regulatory positioning disclosuresA published weekly report on how large participants are positioned. What you see is a presentation of a public filing, not an estimate of one.
Country macroeconomic indicatorsOfficial releases carried with the country and period they belong to, so a driver traces back to the release it came from.
National and regional statistical officesLocal official series for markets the international agencies do not cover in depth.

Quality gates, traceability, repeatability

  • You can check the inputs, not just the outputs. Every value carries its source and the time it was loaded, so any number on any screen can be traced back to where it came from and when. That is what makes the rest of this page auditable rather than merely asserted.
  • Implausible data is held back, not smoothed over. A plausible-looking wrong number does more damage than an obvious gap, so questionable values are quarantined rather than quietly averaged away. The mechanics are covered on How Market Lens Works.
  • Re-reading history gives the same answer. An update never double-counts, so the record you audit today is the record you audited last month plus whatever has genuinely happened since — which is what lets you hold us to a figure we published a year ago.

The transparent-ML promise

Two commitments, and they are the reason this page exists at all. See the track record first: the Scorecard is there so that you judge a model's reliability before you rely on its forward number, rather than after a budget round has already been built on it. Honest uncertainty: confidence bands are drawn from real, measured volatility. They widen as the horizon lengthens and as a market turns volatile, and they tighten when it is calm, so the range you plan against tracks the conditions you are actually in — and the far horizons, where a plan is hardest to reopen, are given the width they deserve.

It is worth being equally clear about scope. Market Lens does not assess prices and is not a settlement reference: keep the assessed benchmarks your contracts already reference, and use Market Lens as the decision layer on top of them — the forward view, the range around it, and the published record of how that view has performed.

AI-generated advisory reports and news-sentiment summaries are decision support, not financial advice. They are strong at rapid synthesis and context, and they are commentary on the numbers rather than additional evidence beyond them. Confirm anything material against the underlying data and your own judgment before you commit.

You now know how to read the record. The next question is what produced it — the sources, the pipeline and the backtest protocol behind every score on this page.

Read the methodology