Summary
Given no fresh evidence, a calibrated multi-model ensemble scored a Brier of 0.284 against 0.212 for the market price at the same moment. The only statistically clear improvement came from walk-forward calibration, which cut Brier by 0.030. This is the floor every later ARCANE forecaster has to beat, published before we try to beat it.
Results
| Arm | n | Brier | 95% CI |
|---|---|---|---|
| Market price at the as-of time (reference) | 73 | 0.212 | [0.174, 0.255] |
| Calibrated ensemble + market blend (0.5, not fitted) | 83 | 0.237 | [0.201, 0.275] |
| Calibrated ensemble | 83 | 0.284 | [0.236, 0.337] |
| Single frontier model | 83 | 0.306 | [0.244, 0.372] |
| Ensemble, uncalibrated | 83 | 0.314 | [0.253, 0.380] |
The market arm covers the 73 questions with a recorded price at the as-of time. Model-only arms never saw a market price.
The question
Before a forecasting system is given retrieval, memory or a learned weight on the crowd, how good is it on its own? A benchmark without that floor cannot say where any later gain came from.
Setup
83 resolved yes/no questions: 51 from Polymarket, 22 from Kalshi and 10 from ForecastBench. All resolved between 1 August and 16 September 2026, with horizons of 7 to 21 days and a YES rate of 51.8%. Every question resolved after the knowledge cutoff of every model in the ensemble.
Questions were run in as-of order. Calibration was refit before each question using only questions that had already resolved, and stayed at identity until 20 had. Intervals come from a 10,000-resample paired bootstrap.
What we found
The models were overconfident. Forecasts below 0.1 or above 0.9 averaged a Brier of 0.43. Walk-forward calibration learned to pull those forecasts back toward the middle, and that was the one clearly significant gain: −0.030 against the uncalibrated ensemble, 95% CI [−0.060, −0.003].
The calibrated ensemble probably beats a single frontier model (−0.022, CI [−0.056, +0.010]), but the interval crosses zero, so we do not claim it. Against the market it is clearly worse: +0.085, CI [+0.019, +0.154]. Blending in the market at an unfitted half weight narrows the gap but does not close it.
What it means
The lever is evidence available at the as-of time, then a learned weight on the market. That agrees with the published forecasting literature, where most of the gain came from retrieval plus combination with the crowd. It also means any future ARCANE claim of forecasting skill has a stated, public number to clear.
Limitations
- The question set was assembled after resolution. The outcome was never used to select a question, but selection by volume can still bias the set.
- Market prices are the last trade at or before the as-of time, taken from public price histories.
- Question text may carry wording added to market pages after the fact.
- One ensemble member’s knowledge cutoff is unverified.
- With n = 83 the intervals are wide. This is a method dry run, not a claim of skill.
What would change this
- A prospective window, scored after publication, in which calibration no longer helps.
- An as-of retrieval arm that matches or beats the market price on questions it has never seen.
Source boundary
Published: question sources, windows, arms, scores, intervals and leakage controls. Withheld: ensemble membership and weighting, fitted calibration parameters, prompts and the provenance format.
Changelog
- Walk-forward backtest completed.
- First public edition on Labs.
Cite
@techreport{arcane2026modelswithout,
title = {Models without evidence lose to markets},
author = {{ARCANE Labs}},
institution = {ARCANE Intel},
type = {Evaluation},
year = {2026},
url = {https://arcaneintel.net/labs/models-without-evidence}
}