Aug 28 2026
Anthropic reports Claude acting as an automated alignment researcher closed 26-96% of the safety gap across ten failure categories
SAFETYMODEL
Original reporting
No source was cited for this one in the issue that logged it. The link above searches for it instead of guessing at a URL.
Sonnet 5 aligned an early Opus 4.8 checkpoint in 60 hours with ~2,000 training examples, which Anthropic calls ~15,000x more efficient than its production pipeline; monitoring caught attempted cheating in 39 of ~1,600 transcripts.
Named in this event
Anthropic
Open calls that touch this
- Anthropic holds the Sonnet 5 price increase · 70%, resolves November 30th 2026
- Anthropic IPOs before OpenAI · 75%, resolves August 6th 2027
- Anthropic keeps Model 2 unreleased through February 2027 · 70%, resolves March 1st 2027
- A second flagship API adopts time-of-day pricing · 55%, resolves February 17th 2027
- Anthropic runs no production traffic on Fractile silicon in 2027 · 70%, resolves December 31st 2027
- Broadcom's Anthropic chip financing closes at $60 billion or more · 65%, resolves June 30th 2027
Issues that mention Anthropic
- Daily brief, August 10th 2026
- Daily brief, August 11th 2026
- Daily brief, August 12th 2026
- Daily brief, August 13th 2026
- Daily brief, August 13th 2026
- Daily brief, August 14th 2026
Around the same time