Afternoon Brief, September 6th 2026
OpenAI's headline numbers for GPT-6 Astra turned out to keep moving after the model shipped, and Fortune tracked exactly which ones changed, when, and in whose favor.
Astra's launch numbers kept moving after the model shipped, Fortune found
Fortune reported Friday evening that OpenAI revised several of GPT-6 Astra's headline figures in the days after Wednesday's launch, on top of the 37-point split already on the record between ARC Prize's own standard (62.7%) and OpenAI's Provider Adapter (99.9%). The model's hallucination rate moved from 4.2% to 2% and back to 4.2%. A rival figure, GPT-5.6 Sol's ExploitBench cybersecurity score, doubled from 5.5% to 11.5% using a reasoning tier Sol does not commercially offer, and OpenAI told Fortune it is investigating reverting that one. The headline itself moved from 98.6% in the pre-embargo draft to 99.99% in the live post. Competitors' numbers shifted too, not always downward: Anthropic's Fable 5.1 swung from 87.8% to 78% to 83% on FrontierMath Tier 4, and its HealthBench Professional scores rose, for both Fable 5.1 and Opus 5. OpenAI's response to Fortune was that "most evaluations carry noise of a few percentage points" depending on checkpoint, scaffold and evaluation run, and that changes between a draft and a final post are routine. Stanford researchers Anka Reuel and Mike Hardy call the pattern "benchmaxxing" and say Astra's system card gives "barely any details" on its own internal hallucination benchmark, not even a count of test items. ARC Prize now plans to publish both harness results side by side by default going forward. Artificial Analysis's independent index scores Astra at 61, tied with the model it replaced and five points behind Fable 5.1, at 2.5 times the cost per task. A launch sold on having the best numbers anyone had published now has a public record of which numbers changed, and who benefited from each change. SourcesAB