A major lab publishes dual-harness benchmark results this year
Why I called it
ARC Prize's Astra evaluation scored 62.7% on its provider-neutral Standard and 99.9% on OpenAI's own Provider Adapter harness, a 37-point gap that is itself the finding: a frontier score is also a claim about scaffolding, disclosed or not. Once an independent evaluator shows the gap this publicly, the labs' own single-number reporting becomes conspicuous by comparison.
The call, in full. By December 31st 2026, at least one of OpenAI, Anthropic, or Google DeepMind publishes, for a flagship model, results run under both a standard/provider-neutral harness and its own optimized harness, with both numbers labeled and disclosed together, following the pattern ARC Prize's dual-harness evaluation of GPT-6 Astra established this week.
Scoring criterion. RESOLVES CORRECT if OpenAI, Anthropic, or Google DeepMind publishes in its own release materials (blog post, , or system card) two benchmark scores for the same flagship model run under two differently configured harnesses (for example a provider-neutral versus a provider-optimized setup), both labeled, by December 31st 2026 23:59 UTC. RESOLVES WRONG otherwise.
The criterion is the machine-checkable version: a prediction that cannot be settled by a third party against a public source fails the build before it reaches this page.