Platform
How Market Lens Works: Data, Models & Backtesting
You are being asked to put a number from somebody else's model into your budget, or to size a position against it. For that to be reasonable, three questions need answering first: where did the input data come from, what was done to it on the way to becoming a forecast, and how does anyone know whether the output has ever been right? This page answers all three, in that order. It is deliberately about process rather than machinery — what gets measured, what gets thrown away, what gets tested, and what is not claimed — because process is the part you can actually inspect, and process is what determines whether a forecast is worth planning against.
One thing this page does not contain is an accuracy figure. That is on purpose, and the reasoning is in the backtest section below.
Where the data comes from
Market Lens does not run on a single feed. Prices, drivers and positioning arrive from independent classes of source, and the breadth is structural rather than decorative: when two independent sources describe the same market and disagree, that disagreement is itself a signal to check something before it reaches a screen. The classes below are described by what each contributes, because what matters to a forecast is the character of an input — how fast it updates, how far back it goes, how consistently it is compiled — rather than whose name is on it.
| Class of source | What it contributes |
|---|---|
| International statistical agencies | Long, methodologically consistent price and activity series compiled by public bodies. Slower to publish than a market feed and far more comparable across years and countries, which is what makes them the right backbone for anything with a multi-year horizon. |
| Market and futures price feeds | Daily traded prices, plus the forward curve of contract prices across delivery dates. This is the fast-moving layer: it is what makes a daily forecast possible at all, and it is the yardstick every forecast is ultimately scored against. |
| Regulatory positioning disclosures | The published weekly breakdown of how large participants are positioned, split between commercial hedgers and speculators. One of the few genuinely forward-looking inputs available anywhere, because it describes intent rather than outcome. |
| Country macroeconomic indicators | Demand-side drivers — activity, labour, prices, confidence, trade, housing — across many countries, so that a commodity's own price history is never the only thing a model is allowed to see. |
| National and regional statistical offices | Local price series for markets the international agencies do not cover in depth. Unglamorous, and often where the most useful series for one specific supply chain turns out to live. |
One concession, stated plainly rather than buried: contract-grade assessed benchmarks are not part of this — by design. Market Lens does not assess a price and has no ambition to be the reference your contracts settle against. Keep the assessed benchmarks your contracts already reference; Market Lens is the decision layer on top of them — the part that tells you where that benchmark is likely to go, and how much room to leave for being wrong.
From raw data to a published forecast
- Ingestion, with provenance attached — Each source is pulled on its own rhythm and stored as it arrived, unmodified, with its origin and load time recorded against every value. Keeping the raw arrival separate from everything done to it afterwards is what makes any number on any screen traceable back to where it came from — and what makes a correction possible without quietly rewriting history.
- Cleaning and quality checks — Incoming data then passes automated checks aimed at the things that ruin a model silently: zero or missing prices, impossible spikes, values that contradict the rest of their own series. Failures are flagged and quarantined rather than smoothed over, because a plausible-looking wrong number does far more damage downstream than a visible gap. Updates are repeatable by construction — re-running one never duplicates or double-counts a record — so the history you read today is the history you read last week plus what has genuinely happened since.
- Feature construction — The model is never shown a bare price. It is shown the commodity's own behaviour across several time scales, the behaviour of the markets that historically move with it, how large participants are positioned, and the macroeconomic drivers relevant to that market. Most of the real work in forecasting sits in this step, because a forecast is largely a function of what you decided the model was allowed to look at.
- Fitting, one commodity at a time — A statistical model is then fitted to each commodity's own price history and the drivers that move it — separately, per commodity and per grain — grain being the monthly or daily view you read it in. There is no single global model asked to be equally good at natural gas and at cocoa. A market's own history is the best available guide to its own behaviour, and fitting per market is what lets a mean-reverting commodity be handled differently from a trending one without anyone hand-tuning a narrative. Fits are refreshed as new data arrives.
- Publication, with a version stamp — What reaches the product is the model's output, published as produced and stamped with the version of the model that produced it — and that version is visible on screen next to the numbers. So if a newer model changes the story for a commodity, the change is attributable to a version you can point at, rather than surfacing as an unexplained shift in next month's plan.
Everything downstream of that chain inherits it. The forecast pages, the regime and risk surfaces, the correlation and hedge screens and the scenario simulator all read the same cleaned series, so no two screens are working from different versions of the same price history — see Price Forecasts: Monthly & Daily for how the published output is laid out and read.
The backtest discipline
A model fitted to a full history can nearly always be made to look good on that history. Judging it that way teaches you nothing you need to know, because the question in front of you is not how well a model describes a past it was shown — it is how well it would have called a future it had not seen. A backtest is the procedure that answers the second question, and it is the only test whose shape matches the way you actually use a forecast.
- Cut the history at a date in the past — Choose a cut-off and treat everything after it as though it has not happened yet.
- Re-fit on only what was known then — The model is fitted again from scratch on data up to that cut-off and nothing beyond it. This is the step that is easy to skip and fatal to skip: a model that has quietly been shown the answer will score beautifully and tell you nothing about tomorrow.
- Forecast forward, then freeze it — It produces a forecast across the same horizons the live product publishes, and that forecast is frozen. No revisions, no second attempt once the outcome is known.
- Score the frozen forecast against what actually happened — The prices that followed the cut-off are then revealed and compared against the frozen forecast, horizon by horizon. This is out-of-sample scoring in the strict sense: at the moment that forecast was made, none of the data it is now being judged against existed as far as the model was concerned.
- Roll the cut-off forward and repeat — five windows in all — A single window can flatter or punish a model by luck; one calm quarter proves nothing about a turbulent one. Repeating the whole procedure across five rolling windows is what turns an anecdote into a track record, and it is what makes it possible to report a result per commodity and per horizon instead of one comforting number for the platform as a whole.
That is the whole argument for the discipline: it is the only procedure that answers the question you really have. Not “is this model clever”, but “would this forecast have helped me at the time, on the commodity I buy, at the lead time I plan against”. It is also unforgiving in a useful direction — a model that only looks good with hindsight has nowhere to hide from it — which is exactly the property you want in the test that decides how much weight a number deserves.
This page states no accuracy figure — not an average, not a range, not a best case. An average across every commodity and horizon would flatter some markets, libel others, and still not be the number you need. The number you need is the record for the commodity you buy, at the horizon you plan against, and it is published per commodity and per horizon on Forecast Accuracy & Why You Can Trust the Data. Go and read the one that applies to you.
Why the confidence band widens
Every published forecast is a band as well as a line, and the band is where the honesty lives. It is drawn as a 90% range, and it is time-varying: it widens as the horizon lengthens, widens further when the market itself is volatile, and tightens when a market is calm. None of that is a stylistic choice. Uncertainty about a price genuinely compounds with time — tomorrow's price is tightly constrained by today's, next winter's is constrained by very little — and volatility is measurable, so a band that ignored either would be understating risk in every number it drew.
Which is the plain version of the point: a band that does not widen with horizon is a band that is lying to you. A fixed-width margin is easier to draw and it fails in exactly one direction — it understates the far horizon, which is precisely where a budget round or a hedge is most exposed and least recoverable. Widening bands cost a forecast some visual confidence and buy you a plan that survives contact with the market. That trade is worth making every time.
| What the band looks like | What it is telling you |
|---|---|
| Narrow, and staying narrow across the horizon | A market the model can pin down. You can plan close to the central path and keep the cushion modest. |
| Wide from the very first period | A volatile market. Hold more cushion, or stagger commitments instead of committing all at once — the near term is already uncertain here. |
| Narrow near term, fanning out quickly | The next few periods are knowable and the far ones are not. Budget tightly for the near months and treat the far ones as a range to survive rather than a number to commit to. |
The band makes a claim, so the claim is checked. A 90% band should contain the price that actually happened about 90% of the time; whether it does is measured out of sample and published per commodity as calibration coverage. A band that holds its promise means the cushion you size from it is right-sized rather than wishful — which is why calibration, not the headline error, is usually the more useful column on the scorecard. Every measure named here is defined in plain language in the Analytics Glossary.
How the risk metrics are produced
A forecast tells you where a price is likely to go. The risk analytics answer a blunter question: if it goes against me, how much? Both headline figures are computed from the same raw material as everything else — the commodity's own observed price moves, and how turbulent conditions are right now — rather than from anybody's opinion about how bad things could get.
- Value-at-Risk is the move you would expect not to exceed except in the worst cases — roughly one period in twenty at the 95% level, one in a hundred at 99%, over whichever window the figure is quoted for. Read it as the size of the cushion a normal bad stretch calls for.
- Conditional Value-at-Risk, also called Expected Shortfall, is the average loss across the cases that do break through that line. It answers the follow-up question — when the cushion is breached, how deep does the hole actually go — and it is always at least as large as Value-at-Risk, which makes it the more conservative figure to plan against.
Each figure is produced in more than one way on purpose, because each way makes a different assumption and the differences are the interesting part. One is a clean statistical baseline that assumes ordinary market behaviour. One is read straight from the commodity's real past moves, trusting observed history over tidy assumptions. One reacts quickly to recent shocks, so it rises fast when a market turns rough instead of averaging the turbulence away. And one is scaled by how stressed the broad market is, so a nervous market widens the estimate for everything at once. When those readings agree, the risk picture is stable. When they diverge, the gap is the finding: recent conditions are saying something the long-run average is not. The surfaces that present all of this, alongside what kind of market a commodity is currently in, are covered in Market Regimes & Risk Analytics.
For a buyer this converts directly into money. A worst-case percentage move applied to your annual spend on a commodity is the cushion that exposure deserves, and the conditional figure tells you what becomes of that cushion in a genuinely bad month rather than a merely disappointing one. The practical value is not the sophistication — it is that a cushion sized from a measured distribution of real past moves is a number you can defend in an approval meeting, and a round number that felt about right is not.
Market intelligence and analytics — not investment advice. A forecast is a modelled range of outcomes, not a promise about a price, and every model here is a simplification of a market under no obligation to keep behaving as it has. Use these numbers to size cushions, rank exposures and time decisions, then confirm anything material against the underlying data and your own judgment.
This page tells you how the numbers are made. The scorecard tells you how they have actually done — commodity by commodity, at the horizon you plan against.
Inspect the track record