No. 05

Evaluation

Models without evidence lose to markets

The walk-forward baseline for the ARCANE forecasting benchmark: 83 resolved questions, and the one gain that held.

Result
0.212 vs 0.284
Outcome
Negative result
Updated
25 Sep 2026
Fig. 1 · Drawn from the published numbers
0.212 · THE NUMBER TO BEATMARKET PRICEN 73 · REFERENCE0.212ENSEMBLE + MARKET BLENDN 830.237CALIBRATED ENSEMBLEN 830.284SINGLE FRONTIER MODELN 830.306ENSEMBLE, UNCALIBRATEDN 830.314+0.085 AGAINST THE MARKET0.150.200.250.300.350.40BRIER SCORE · 95% INTERVAL · LOWER IS BETTER83 QUESTIONS · 51 POLYMARKET · 22 KALSHI · 10 FORECASTBENCH0.212 · THE NUMBER TO BEATMARKET PRICE0.212ENSEMBLE + MARKET BLEND0.237CALIBRATED ENSEMBLE0.284+0.085 AGAINST THE MARKETSINGLE FRONTIER MODEL0.306ENSEMBLE, UNCALIBRATED0.3140.150.200.250.300.350.40BRIER · LOWER IS BETTER
Fig. 1Brier score by arm with its 95% bootstrap interval, on 83 resolved yes/no questions; lower is better. The dashed line is the market price at the as-of time. Red is the calibrated ensemble, 0.085 worse than the market, interval [+0.019, +0.154]. Each short line at the foot is one question, grouped by source.

Summary

0.212 vs 0.284Brier score, market vs calibrated ensemble (lower is better)

Given no fresh evidence, a calibrated multi-model ensemble scored a Brier of 0.284 against 0.212 for the market price at the same moment. The only statistically clear improvement came from walk-forward calibration, which cut Brier by 0.030. This is the floor every later ARCANE forecaster has to beat, published before we try to beat it.

Results

Brier score by arm, with 95% bootstrap intervals
ArmnBrier95% CI
Market price at the as-of time (reference)730.212[0.174, 0.255]
Calibrated ensemble + market blend (0.5, not fitted)830.237[0.201, 0.275]
Calibrated ensemble830.284[0.236, 0.337]
Single frontier model830.306[0.244, 0.372]
Ensemble, uncalibrated830.314[0.253, 0.380]

The market arm covers the 73 questions with a recorded price at the as-of time. Model-only arms never saw a market price.

The question

Before a forecasting system is given retrieval, memory or a learned weight on the crowd, how good is it on its own? A benchmark without that floor cannot say where any later gain came from.

Setup

83 resolved yes/no questions: 51 from Polymarket, 22 from Kalshi and 10 from ForecastBench. All resolved between 1 August and 16 September 2026, with horizons of 7 to 21 days and a YES rate of 51.8%. Every question resolved after the knowledge cutoff of every model in the ensemble.

Questions were run in as-of order. Calibration was refit before each question using only questions that had already resolved, and stayed at identity until 20 had. Intervals come from a 10,000-resample paired bootstrap.

What we found

The models were overconfident. Forecasts below 0.1 or above 0.9 averaged a Brier of 0.43. Walk-forward calibration learned to pull those forecasts back toward the middle, and that was the one clearly significant gain: −0.030 against the uncalibrated ensemble, 95% CI [−0.060, −0.003].

The calibrated ensemble probably beats a single frontier model (−0.022, CI [−0.056, +0.010]), but the interval crosses zero, so we do not claim it. Against the market it is clearly worse: +0.085, CI [+0.019, +0.154]. Blending in the market at an unfitted half weight narrows the gap but does not close it.

What it means

The lever is evidence available at the as-of time, then a learned weight on the market. That agrees with the published forecasting literature, where most of the gain came from retrieval plus combination with the crowd. It also means any future ARCANE claim of forecasting skill has a stated, public number to clear.

Limitations

  • The question set was assembled after resolution. The outcome was never used to select a question, but selection by volume can still bias the set.
  • Market prices are the last trade at or before the as-of time, taken from public price histories.
  • Question text may carry wording added to market pages after the fact.
  • One ensemble member’s knowledge cutoff is unverified.
  • With n = 83 the intervals are wide. This is a method dry run, not a claim of skill.

What would change this

Any one of these would change the conclusion

  1. A prospective window, scored after publication, in which calibration no longer helps.
  2. An as-of retrieval arm that matches or beats the market price on questions it has never seen.

Source boundary

Published
question sources, windows, arms, scores, intervals and leakage controls.
Withheld
ensemble membership and weighting, fitted calibration parameters, prompts and the provenance format.

Changelog

  1. Walk-forward backtest completed.
  2. First public edition on Labs.
  3. A figure drawn from the published numbers added. No wording or number changed.

Cite

@techreport{arcane2026modelswithout,
  title       = {Models without evidence lose to markets},
  author      = {{ARCANE Research}},
  institution = {ARCANE Intel},
  type        = {Evaluation},
  year        = {2026},
  url         = {https://arcaneintel.net/research/models-without-evidence}
}