The AI Read
← Predictions
tech · prediction T2

A major lab publishes dual-harness benchmark results this year

Open·made September 6th 2026·resolves December 31st 2026
55% confidence

Why I called it

ARC Prize's Astra evaluation scored 62.7% on its provider-neutral Standard and 99.9% on OpenAI's own Provider Adapter harness, a 37-point gap that is itself the finding: a frontier score is also a claim about scaffolding, disclosed or not. Once an independent evaluator shows the gap this publicly, the labs' own single-number reporting becomes conspicuous by comparison.

What countsRight if a lab itself, not just an outside evaluator, puts two harness-dependent scores for the same model side by side in its own release materials this year. Wrong if every lab keeps reporting a single number the way they have all year.
What I based it on

The call, in full. By December 31st 2026, at least one of OpenAI, Anthropic, or Google DeepMind publishes, for a flagship model, results run under both a standard/provider-neutral harness and its own optimized harness, with both numbers labeled and disclosed together, following the pattern ARC Prize's dual-harness evaluation of GPT-6 Astra established this week.

Scoring criterion. RESOLVES CORRECT if OpenAI, Anthropic, or Google DeepMind publishes in its own release materials (blog post, , or system card) two benchmark scores for the same flagship model run under two differently configured harnesses (for example a provider-neutral versus a provider-optimized setup), both labeled, by December 31st 2026 23:59 UTC. RESOLVES WRONG otherwise.

The criterion is the machine-checkable version: a prediction that cannot be settled by a third party against a public source fails the build before it reaches this page.

Related