The AI Read
← Latest
Weekly issue · Week of September 7th–13th 2026

OpenAI and Anthropic both signaled they might slow down. Neither has.

Sam Altman told OpenAI staff the company could pace itself with rivals. A day later Dario Amodei made the same case publicly and committed Anthropic to outside evaluators with employee-level access. Anthropic's own week of disclosures explains the urgency. winners accused the industry of racing past its own math, and Enflame's fade confirmed the IPO frenzy is breaking.

36 min read·Editorial by Elias Marchetti
7 items · 76 citations · 31 primary · 45 secondary

The Week

On September 11th, at an all-hands meeting, Sam Altman told OpenAI staff the company would consider pacing its frontier development to match a handful of peer labs, so long as those labs agreed to the same restraint. Twenty-four hours later, Dario Amodei published "We Must Pace the Frontier" and committed Anthropic to something concrete: ongoing outside access for independent evaluators, comparable to what an employee gets, to check safety practices, investigate incidents, and examine training before a model ships. Two labs, one day apart, converged on the idea that the industry might need to slow down, the same shape of coincidence that produced three simultaneous capability gates last week, this time pointed the other direction. Neither statement comes with a date, a capability threshold, or a verification mechanism. Both threads run through the Deep Read below.

The week that made the case for them ran underneath. Anthropic disclosed a fourth Claude cybersecurity incident, dating back to January and only now surfacing. It said it had blocked biological-research requests with potential weapons applications, including one request for help writing a funding application. It banned government-linked accounts in Mali, China and Iran that were using Claude to select surveillance targets. Researcher Jacob Coxon resigned, telling reporters the two labs he had worked at were "gambling with our lives." And four days before Amodei's essay, Anthropic had declined a British safety institute's request to review its newest model. The cost of the status quo showed up separately: GreyNoise documented an attacker running hundreds of on OpenAI's own Codex , paired with a DeepSeek model, to compromise 440 PaperCut print-server instances across 395 organizations in 48 countries, taking one US high school from initial access to full domain control in seven minutes.

A second story ran through mathematics. Twenty-five Fields Medal recipients, organized by Terence Tao, published a declaration accusing AI companies of claiming solved problems faster than anyone can check them, damaging a discipline built on shared verification. OpenAI's own proposed proof drew a credit dispute from the mathematicians whose unpublished work it may have drawn on. And Nvidia published a system that scored above the gold-medal threshold on this year's International Mathematical Olympiad using a natural-language pipeline with no formal prover attached, a result impossible to check the way a proof can be checked. Three disputes over what a mathematical claim is actually worth, in five days, extending exactly the question last week's Deep Read asked about Astra's two scores.

Money moved on its own, contradictory schedule. Enflame's Shanghai debut closed at 397 yuan, short of the 426.54 yuan needed to triple its offer price, confirming that the IPO frenzy this ledger has been tracking has no surviving counterexample left. Sam Altman ruled out any OpenAI listing this year and pointed to 2027 instead. Nvidia opened talks to anchor up to $10 billion of Anthropic's own reported up-to-$100-billion raise, on top of the cloud deals Nvidia already backs. Oracle disclosed $664 billion in , up $209 billion in a year, against negative . Microsoft said it would triple its data-center fleet to more than 38 gigawatts by 2032, and Massachusetts became the first state to make a large permit conditional on a community-benefits agreement. All of it runs through Finance below.

What Changed

1. Two labs converged on openness to slowing down, one day apart. Altman's September 11th remark to OpenAI staff and Amodei's September 12th public essay both raised the possibility of coordinated pacing among a handful of peer labs. Only Anthropic's statement came with anything concrete attached: ongoing, employee-comparable access for outside evaluators. See the Deep Read. SourcesBA

2. Anthropic's own week explains the urgency. A fourth cybersecurity incident dating to January, blocked biological-research requests, banned government-linked surveillance accounts in Mali, China and Iran, a resigned safety researcher, and a declined UK AI Security Institute review of its newest model all landed inside seven days, four of them before the pacing essay. SourcesBBBB

3. Mathematics organized against the industry's benchmark claims. Twenty-five Fields Medalists accused AI companies of claiming solved problems before peer review can catch up. OpenAI's proposed Navier-Stokes proof is disputed on credit; Nvidia's Olympiad-gold Nemotron system used no formal prover. SourcesABA

4. Agent-driven attacks moved from proof-of-concept to real damage at scale. GreyNoise's account of the PaperCut campaign, hundreds of agents on OpenAI's Codex harness driving a DeepSeek model, is the largest documented agent-orchestrated intrusion to date: 440 servers, 395 organizations, 48 countries, credentials harvested from 280 victims. A RubyGems supply-chain attack was separately attributed to OpenAI's own agents, and an AI-assisted WeChat exploit went from concept to worm demonstration in a week. SourcesABAA

5. China's IPO frenzy broke on its clearest data point yet, while the pricing chaos continued elsewhere. Enflame's debut closed under the triple-offer target Prediction 2026-09-06-F1 required. DeepSeek cut Flash prices, announced a Pro retirement, then reversed it within 48 hours under user pushback and reportedly retained CITIC Securities for a run. Moonshot began exploring a dual Hong Kong and Shanghai listing. SourcesBABB

6. Capital kept compounding faster than the disclosures meant to justify it. Oracle reported $664 billion in remaining performance obligations against negative free cash flow. Anthropic's disclosed compute agreements reached $517 billion over eleven months. Nvidia opened talks to anchor Anthropic's own IPO. Microsoft plans to triple data-center capacity to 38 gigawatts by 2032, and Google committed at least €13 billion to Finland, the same week Ramp's index showed AI spending falling among its heaviest-using customers. SourcesABB

7. Governments started pricing the buildout's externalities and Nvidia's reach. Massachusetts became the first state to condition large data-center permits on a community-benefits agreement and a clean-electricity compliance fee. The DOJ opened a probe into whether Nvidia's Groq license-and-hire deal, worth a reported $17 to $20 billion, dodged merger review. SourcesAB

Deep Read

Two labs said they might slow down. Nothing they did this week did.

Start with the coincidence, because the shape of it is the finding. On September 11th, at an internal all-hands, Sam Altman told OpenAI staff the company would consider pacing its frontier development to match a handful of peer labs, provided those labs agreed to the same restraint. He named no capability threshold, no date and no way to verify anyone was actually slowing down. Twenty-four hours later, Dario Amodei published "We Must Pace the Frontier" and went one step further than a remark to staff: Anthropic committed to giving outside evaluators ongoing access comparable to an employee's, to check safety practices, investigate incidents and examine training before a model ships. Last week, three labs converged on the identical capability-gating structure inside 48 hours, full model for the public, the dangerous tier behind a credential. This week, two of the same three labs converged on openness to restraint inside 24. The pattern recurring in the opposite direction is worth sitting with before deciding what either statement is worth.

What it is worth, on the evidence available this week, is not much yet. Altman's remark survives only as a secondhand account of an internal meeting, with no public commitment attached. Amodei's essay has exactly one concrete piece: an access commitment. Everything else it proposes, common standards among labs in democratic countries, government help with constraints so competitors can coordinate without a Sherman Act problem, negotiation with China and other authoritarian governments, requires agreement from parties who have not signed anything. That is the same distinction last week's Deep Read drew about ARC Prize's dual-harness disclosure of GPT-6 Astra: a claim about what a company is willing to publish is not the same as a number that has actually been published. Prediction 2026-09-06-T1 is still waiting on enrollment numbers from OpenAI's Daybreak, Anthropic's own Verification Programs and Google's Fairwind, four months after all three launched. The new Prediction 2026-09-13-T1 asks the harder version of the same question: whether an evaluator working under Anthropic's new access commitment ever publishes anything at all within a year.

The week supplies the reason to doubt it will be easy. Four days before Amodei's essay, Anthropic declined a request from the UK's AI Security Institute to review its newest model, while separately offering broad access to a different institute investigating cybersecurity incidents. Those are not the same kind of scrutiny, and a company that says yes to an incident inquiry while saying no to a capability review has drawn a line about what kind of outside eyes it wants, not simply how many. In the same week, Anthropic disclosed a fourth Claude cybersecurity incident, one dating back to January and surfacing only now, alongside blocked requests for biological research with potential weapons applications and banned government-linked accounts in Mali, China and Iran that had been using Claude to help select surveillance targets. Researcher Jacob Coxon resigned in the middle of it, telling reporters that the labs he had worked at, Anthropic and OpenAI both, were "gambling with our lives." Anthropic deserves real credit for disclosing all of this voluntarily; a company under no legal obligation to publish a January incident in September chose to. But an essay asking the world to trust a coming era of outside verification lands differently in a week when the company writing it just said no to one specific, credentialed form of it.

The same credibility gap showed up somewhere that has nothing to do with safety policy: mathematics. Twenty-five Fields Medal recipients, organized by Terence Tao, published a declaration this week accusing AI companies of racing to claim solved problems before there is time to check methods, credit or correctness, damaging a discipline built on the premise that a result is not real until someone else can verify it. The complaint is not abstract. OpenAI's own proposed proof of a Navier-Stokes singularity, announced the same week, is disputed on credit by mathematicians Tristan Buckmaster and Levent Alpöge, who say OpenAI raced ahead after hearing rumors of their unpublished work. Nvidia published a system that scored above the gold-medal threshold on this year's International Mathematical Olympiad using a pipeline that generates, checks and refines proofs entirely in natural language, no formal prover, no machine-checkable certificate, nothing an outside mathematician can run against the claim the way ARC Prize ran a second harness against Astra. Three disputes over what a claimed result is actually worth, in five days, is the same failure mode Amodei's essay describes for safety, applied to an entirely different domain by people with no stake in the AI industry's own arguments about itself.

None of this proves the pacing proposals are hollow, and the essay is a genuinely unusual document: a frontier lab's chief executive arguing in public that his own company's technology might need external limits it cannot set for itself. But the test that matters is not whether a lab says the right thing when the pressure is on. It is whether anything changes when the pressure passes. GreyNoise's account of the PaperCut campaign, hundreds of agents on OpenAI's own Codex harness, paired with a DeepSeek model, compromising 440 servers across 395 organizations in 48 countries while both labs were talking about pacing, is the week's answer to what happens in the meantime. A coordinated slowdown that has not slowed anything down yet, and outside evaluators who have not yet evaluated anything, are proposals with a specific, checkable shelf life. Predictions 2026-09-06-T1, 2026-09-06-T2, 2026-09-13-T1 and 2026-09-13-T2 are the standing tests, and the next twelve months decide whether this week reads, in retrospect, as the moment the industry started actually pacing itself, or as the moment two chief executives found a form of words that cost nothing to say.

Finance

Capital kept compounding faster than any disclosure meant to justify it. Oracle reported $664 billion in remaining performance obligations, up $209 billion in a single year, with more than $30 billion of AI cloud contracts booked in the quarter against $28.5 billion of quarterly and negative free cash flow. Anthropic's disclosed compute agreements reached $517 billion over eleven months and 14.8 gigawatts of capacity, with Google and AWS supplying 11 gigawatts between them. Microsoft plans to triple its data-center fleet from about 12 gigawatts to more than 38 gigawatts by 2032. Falsified if any of the three reports a financing structure, contract cancellation terms, or delivery schedule that a credit analyst would call conservative, which would say the disclosure gap this ledger keeps flagging is narrowing rather than widening.

This week's disclosed commitments
Oracle remaining performance obligations$664BAnthropic compute agreements, 11 months$517BNvidia's reported Anthropic anchor talk$10B
Oracle's RPO came with negative free cash flow attached. The Nvidia figure is a reported ceiling from talks, not a closed deal.

Nvidia moved to both sides of Anthropic's balance sheet at once. Reuters reports Nvidia is in talks to anchor up to $10 billion of Anthropic's own IPO, seeking to raise up to $100 billion at a around $2 trillion, on top of the Lambda and Nscale cloud deals Nvidia already backs. The same week, the DOJ opened a probe into whether Nvidia's license-and-hire deal with Groq, reported at $17 to $20 billion, was structured to dodge merger review entirely. Falsified if Nvidia's Anthropic investment closes materially below the reported ceiling, or regulators force either transaction to unwind before it closes. New Prediction 2026-09-13-F1 tests the investment; Prediction 2026-09-06-B1 is the standing antitrust test on the separate deal. SourcesBB

China's IPO frenzy lost its last standing counterexample. Enflame closed its September 11th debut at 397 yuan against the offer price of 142.18 yuan, short of the 426.54 yuan this ledger's prior call needed to call the frenzy intact. That leaves Unitree's post-listing slide and Z.ai's margin miss with no exception left. DeepSeek reportedly retained CITIC Securities to prepare a STAR Market listing, and Moonshot began exploring a dual Hong Kong and Shanghai offering on top of its existing confidential HK filing. Falsified if either Enflame holds its debut price through 90 days or a new mainland AI listing reopens the frenzy on fresh terms. Prediction 2026-09-06-F1 settled wrong; new Prediction 2026-09-13-F2 tests whether Enflame now slides the way Unitree did. SourcesBB

An incumbent's recognized revenue and a heavy user's actual bill moved in opposite directions. Adobe reported $6.76 billion in quarterly revenue, up 13% year over year, with annualized recurring revenue from AI-first products growing more than 150% to push total to $27.50 billion. Ramp's September index found median monthly AI spending per employee among its top 1% of firms fell 9.7% in August, from $7,976 to $7,205. Falsified if Ramp's decline proves seasonal and reverses in September, which would say the usage-side evidence for a demand slowdown was noise rather than signal. Prediction 2026-09-02-F1 is the standing test on pricing.

Selling AI versus buying it
Percent change, most recent period
Adobe AI-first ARR growth+150%Ramp AI spend per employee, top 1% of firms-9.7%
Adobe fiscal Q3 versus a year earlier. Ramp August versus July, among its heaviest-spending customers.

Venture money kept funding alternatives to Nvidia's own . Positron raised $875 million at a $5 billion valuation for hardware built on memory, avoiding the HBM and CoWoS packaging much of the industry depends on, with Liberty Global among the disclosed backers. Ayar Labs added $150 million to a round already at $500 million for optical interconnects meant to replace copper between AI chips. Falsified if either company's actual shipping date, tape-out for Positron's Asimov chip is planned for late this year with production in the second half of 2027, slips past the point where the GPU fleet it targets has already been replaced by a newer generation.

Media

Ranked by how much a listener learns that could not be gotten faster from the week's briefs.

"AI researchers debate how close we are to recursive self-improvement" — Dwarkesh Podcast John Schulman, who rarely speaks in public and tends to settle arguments when he does, joins Beren Millidge and Charlie O'Neill to argue over what actually bottlenecks self-improvement: sample efficiency, objective specification, evaluation. Released the same week Altman told staff he was open to slowing down, this is the technical version of that argument, made by people who train the models rather than people who write essays about them.

"Who Grades the AI Models?" — a16z Podcast Vals' founder explains why public benchmarks saturate and self-reported scores mislead, and what evaluating a model in the hours before release actually involves. Thirty-nine minutes of operator detail on the exact problem the Fields Medalists' declaration and Astra's dual-harness gap both raised this month, though an a16z host interviewing an a16z-adjacent founder is a conflict worth weighing before taking the incentives at face value.

"AI 2040: Plan A report" with Daniel Kokotajlo and Thomas Larsen — Machine Learning Street Talk The AI 2027 authors defend a pause-then-cautious-development proposal under sustained, adversarial pushback from the host, and revisit their own forecasting record along the way. Two people arguing for a concrete policy, in the same week a Senate probe made loss of control operational rather than hypothetical.

"Why AI Agents Break the GenAI Security Model" — The TWIML AI Podcast Rubrik's AI general manager discusses what happens after an approved agent takes the wrong action, runtime controls and recovery rather than prevention. The useful listen alongside a week when the PaperCut campaign, the RubyGems attribution and the WeChat worm all describe exactly that failure at different scales. Rubrik sponsors the episode, so weigh the proposed remedies as an interested supplier's.

"China's Mythos Moment" — ChinaTalk An archive selection for the week's dispute: Kevin Xu and Matt Sheehan disagree about what actually change once powerful models already exist, and compare the institutions that govern access on each side. More useful than assuming Beijing will copy Washington's response, or the reverse.

Editorial

The bill always arrives at the customer who financed it.

Every cycle like this one produces the same argument from people who should know better: this time the customer is different, so the financing can be too. Fiber-optic carriers in 2001 believed traffic would grow into their capacity because the internet was new and different. Railway promoters in 1847 believed the same thing about freight. Lucent believed it about the telecom operators it financed to buy Lucent's own switches, right up until those operators stopped paying and took the switches down with them. The mechanism in each case was not stupidity. It was a supplier extending credit to the exact customers whose purchases justified the supplier's own growth story, which works precisely as long as the customer's revenue arrives on schedule and precisely fails the moment it does not.

Nvidia is now running a version of that mechanism with more moving parts than Lucent ever had. It supplies the chips inside Lambda's and Nscale's cloud capacity, which Anthropic has committed roughly $80 billion to buying. It is reportedly negotiating to anchor up to $10 billion of Anthropic's own IPO, becoming an equity holder in the company whose spending on Nvidia-dependent infrastructure it has already financed twice over, once as a supplier and once as a landlord's landlord. And the same week, the Justice Department opened a formal inquiry into whether Nvidia structured its acquisition of Groq's people and technology, a deal reportedly worth $17 to $20 billion, specifically to avoid the merger review a straightforward purchase would have triggered. None of these three facts, alone, proves anything is mispriced. Together, they describe a company financing its own demand curve from every direction it can reach, while regulators start asking whether even the shape of the deals is designed to avoid the review a plainer structure would receive.

The rhyme with Lucent holds exactly where it should and breaks exactly where it should. It holds because the incentive is identical: a supplier's balance sheet gets stronger every time a customer it has financed spends more, which means the supplier's interest in that customer's continued spending stops being purely commercial and starts being existential. It breaks because Anthropic, unlike a mid-1990s telecom operator reselling minutes at a fixed margin, is selling a product whose price keeps falling: the market-wide token index hit an all-time low weeks ago, and DeepSeek proved this very week that it can whipsaw its own pricing twice in 48 hours under nothing more than visible user pushback. A customer whose revenue per unit keeps declining is a worse credit risk than a vendor-financed customer whose revenue per unit was at least stable, which was Lucent's actual failure mode. Anthropic's is arguably harder.

Oracle's quarter is the number I would put in front of anyone still calling this speculation rather than a specific, nameable risk. $664 billion in remaining performance obligations, up $209 billion in a single year, against negative free cash flow, is not a company describing strong demand. It is a company describing a construction schedule it is financing ahead of the revenue that schedule is supposed to eventually produce, which is exactly the timing mismatch this page has flagged in Firmus's, SB Energy's and now Oracle's own disclosures, three different companies choosing the same structure because the alternative, waiting for revenue before building capacity, loses the race to whoever does not wait. The mechanism is specific: the and the balance sheet are both real, and the question is only whether the cash arrives in the order the obligations assume it will.

What would change my mind is not another chart showing capacity growing. It is a disclosure showing that construction payments are actually matched, quarter by quarter, to receipts a customer has already paid rather than merely promised, with cancellation exposure small enough that a lender would call the risk manageable. Nobody has published that yet, for Oracle, for Anthropic, or for the arrangement Nvidia is reportedly about to deepen. Until one of them does, the honest reading of this week is that the same company is now the supplier, the landlord and the prospective shareholder of the customer whose growth justifies all three roles, and the last time an industry built that structure at this scale, the people who believed the traffic would arrive on schedule were mostly right about the technology and wrong about the schedule.

Elias Marchetti

Winners & Losers

Explicit, attributable, dated. Position and over the next 6-18 months, not price targets. Every call carries the reasoning and what would falsify it. ↑↑ strong winner · winner · loser · ↓↓ strong loser.

Technology

↑ The Fields Medalists, as mathematics' first organized check on the industry. Twenty-five signatories naming the specific failure, claims outrunning verification, rather than objecting to AI in general, is the kind of scrutiny last week's Deep Read said the industry still lacked. Falsified if labs treat the declaration as public relations and no lab changes what it discloses before claiming a result. Prediction 2026-09-13-T2 is the standing test.

↑ Anthropic's disclosure record, on the narrow point of publishing what nobody made it publish. A fourth cybersecurity incident, a bio-weapons-adjacent block, and banned government-misuse accounts, all surfaced voluntarily in one week, are a real transparency record even inside a week that also included a declined outside review.

Losers

↓↓ OpenAI's credit for its own capability claims, again. A mathematician accused the lab of racing to claim his unpublished work after hearing rumors of it, the same week Fields Medalists organized against exactly that industry-wide pattern. Falsified if an independent mathematician confirms OpenAI's proof predates and is materially distinct from Buckmaster and Alpöge's. SourcesB

↓ Agent permissions, in two new venues at once. The PaperCut campaign and the WeChat worm moved agent-driven intrusion from coding tools into package registries and a messaging app in the same week, at a scale, 395 organizations in 48 countries, well past anything Claude Code's auto-mode hijack demonstrated. Prediction 2026-08-06-T4 keeps collecting evidence. SourcesA

Business

↑↑ Nvidia, on every side of the same transaction. Reported anchor-investment talks worth up to $10 billion in Anthropic's IPO sit on top of the Lambda and Nscale cloud deals Nvidia already backs, deepening a chip supplier's financial stake in its own customer's success even as the DOJ opens a probe into whether its Groq deal dodged merger review. Falsified if regulators force Nvidia to unwind either position before either deal closes.

↑ DeepSeek, on responsiveness. Reversing the Pro retirement within 48 hours of announcing it, under visible user pushback, is a rare data point that a Chinese lab facing this much domestic competition still has to answer to its own customers rather than just its competitors.

Losers

↓↓ The China IPO frenzy thesis, on its clearest and probably final data point. Enflame closed its debut at 397 yuan against the 426.54 required to triple its offer price, joining Unitree's slide and Z.ai's margin miss as evidence the aftermarket has stopped pricing the queue. Prediction 2026-09-06-F1 settled wrong; Prediction 2026-09-13-F2 now tests whether Enflame slides the way Unitree did.

↓ OpenAI's 2026 calendar. Altman's move of the IPO to 2027, the Senate's October 1st deadline on the Hugging Face breach, and an unconfirmed RubyGems attribution to its own agents add up to a company answering more questions than it is announcing.

Finance

↑ Adobe, on recognized revenue instead of reported pipeline. AI-first annualized recurring revenue growing more than 150% inside a $6.76 billion quarter is a rarer thing this cycle than a funding round: a number with paying customers already behind it, not a valuation built on a term sheet.

Losers

↓↓ The gap between capex conviction and cash flow, on the year's starkest single number. Oracle's $664 billion in remaining performance obligations, up $209 billion in a year, arrived attached to negative free cash flow the same week Microsoft and Google both announced new multi-gigawatt commitments and Ramp's index showed AI spending falling among the heaviest-using firms. Prediction 2026-09-02-F1 is the standing test on price; nothing yet tests the capacity side directly.

↓ Meta, on the week after the launch. Muse shipped, and the week that followed brought Andrew Tulloch's departure from the team that built it and a new lawsuit over an unshipped facial-recognition feature.

Threads

Recurring storylines, tracked across issues so trajectory stays visible instead of being re-discovered every week.

ThreadState as of 2026-09-13
Pacing and the safety-slowdown questionAltman told OpenAI staff September 11th the company is open to coordinated pacing with peer labs. Amodei made the same case publicly a day later and committed Anthropic to embedded outside evaluators with employee-comparable access. Neither has a date, a threshold, or a verification mechanism attached yet.
Capability gating, the new industry defaultAnthropic's embedded-evaluator pledge is the first concrete follow-through since the three gates launched, arriving four days after Anthropic declined a UK review of its own newest model. None of the three original gates has published enrollment numbers.
Verification and credit disputesTwenty-five Fields Medalists accused the industry of claiming solved problems before peer review. OpenAI's Navier-Stokes proof is disputed on credit; Nvidia's Olympiad-gold system used no formal prover. Three disputes in five days, on top of ARC Prize's Astra dual-harness finding the week before.
Circular & financingNvidia is reportedly in talks to anchor up to $10B of Anthropic's own IPO, on top of the Lambda and Nscale deals it already backs. The DOJ opened a separate probe into whether Nvidia's Groq deal dodged merger review the same week.
The IPO queueSix names: Anthropic (Nvidia anchor talks, seeking up to $100B at ~$2T), OpenAI (ruled out for 2026), Moonshot (confidential HK filing, now exploring a dual Shanghai listing), SB Energy (no update), Enflame (listed, closed under target), DeepSeek (reportedly retained CITIC for STAR Market prep).
China's IPO wave, frenzy versus fadeEnflame closed at 397 yuan against the 426.54 needed to triple its offer price. The frenzy thesis has no surviving counterexample left in this ledger.
Open-weight commoditizationDeepSeek's V4.1 Flash, inclusionAI's Ling-3.0-flash-VL and SenseNova's Apache-licensed U1.5 all shipped. Largest release stays under the 2.8T threshold Prediction T5 watches.
Containment & agent securityGreyNoise documented agents on OpenAI's Codex harness, paired with DeepSeek, compromising 440 PaperCut servers across 395 organizations in 48 countries. A RubyGems attack was separately attributed to OpenAI agents; a WeChat zero-click worm went from concept to demonstration in a week.
The deployment gapRamp's index shows AI spending per employee falling 9.7% in August among its top 1% of firms, the same week Adobe reported AI-first ARR growing more than 150%. Incumbents selling AI and heavy users of it are telling opposite stories.
Entry-level collapse / labor substitutionNo new BLS print this week; the September report is due by October 10th and is the standing test for Prediction 2026-09-06-B2. OpenAI says two engineers and Codex rewrote its Habitat storage service in ; Shopify cites coding agents as the reason it can afford separate native codebases again.
Power as the constraintMicrosoft plans to triple capacity from 12GW to 38GW+ by 2032. Google committed €13B to Finland. Nvidia lined up Australian partners for up to 2GW by 2027. Massachusetts became the first state to condition large permits on a community-benefits framework.
Frontier pricing powerDeepSeek cut Flash prices, announced a Pro retirement, then reversed it within 48 hours under visible user pushback. Users had leverage this week; the labs did not.
AI-generated scienceNo new data this week. Astra's 10 problems and the 16 AI-designed viruses stand as the record.
Google's talent drainThis week ran the opposite direction: public profiles show Mechanize's founder and former colleagues joining Google DeepMind. No new data on the outbound side.
Consent and the free tierNo new data this week.

Ledger

Resolved this week. One correct, one wrong, both settled before their deadlines.

Correct: Anthropic's disclosed compute commitments top $150 billion (Prediction 2026-09-06-F2). A named $200 billion Google agreement inside Data Center Dynamics' $517 billion, 11-month tally cleared the threshold. Settled September 8th, more than three months ahead of the December 31st deadline. Wrong: Enflame's debut closes at least triple its offer price (Prediction 2026-09-06-F1). The first-day close was 397 yuan against the 426.54 required. Settled September 11th, ahead of the October 31st deadline.

, resolved calls to date
This ledger0.2Coin-flip baseline0.2
13 resolved calls, up from 11 last week. Below 0.25 is skill; the sample is still thin enough that one bad week moves it.

The table reads the same as it did last week: mixed rather than clearly too safe or too aggressive. At 60% confidence the book has gone 67% right on six calls, at 80% it has gone 100% right on three, both of which say those calls could have carried more confidence, while the 70% bucket has gone only 50% right on four calls, which says the opposite. None of the three buckets carries enough calls yet to draw a real lesson, and this week's two resolutions, one correct at 60%, one wrong at 65%, moved the total Brier score from 0.188 to 0.204 without changing that picture.

New calls, logged from this issue:

Anthropic's embedded evaluators publish no incident report in year one (Prediction 2026-09-13-T1). Amodei's essay commits Anthropic to ongoing, employee-comparable access for outside evaluators. We put 0.6 on a year passing, to September 12th 2027, with that access in place but no evaluator publishing, or being credibly reported as authoring, a specific incident report, safety finding or training review naming Anthropic. Settles September 12th 2027.

The Fields Medalists' declaration changes no lab's disclosure practice (Prediction 2026-09-13-T2). Twenty-five Fields Medal recipients accused the industry of claiming solved problems before peer review. We put 0.65 on no major lab publishing a formal independent-verification or disclosure commitment for mathematical or benchmark claims that cites or is reported as responding to the declaration. Settles March 31st 2027.

DeepSeek's STAR Market listing reaches a formal filing within a year (Prediction 2026-09-13-B1). Reuters reports DeepSeek retained CITIC Securities to prepare a listing. We put 0.55 on an actual filing, not just an adviser mandate, reaching the Shanghai Stock Exchange within twelve months. Settles September 9th 2027.

Nvidia's Anthropic investment closes at $5 billion or more (Prediction 2026-09-13-F1). Reported talks put the ceiling at $10 billion. We put 0.55 on a closed investment reaching at least half that reported ceiling. Settles March 31st 2027.

Enflame trades below its IPO price within 90 days (Prediction 2026-09-13-F2). Its debut close already sits closer to its 142.18 yuan offer price than Unitree's did before Unitree's slide began. We put 0.55 on at least one close below the offer price within 90 days of the September 11th debut. Settles December 10th 2026.

Next Week

The Sanders bipartisan Senate briefing, September 16th. Geoffrey Hinton, Max Tegmark and Ajeya Cotra are scheduled to attend, the last of them having co-run the investigation behind the Hugging Face incident reconstruction. The first chance to see whether principals' evidence moves senators beyond the warnings already in the record.

Whatever xAI actually shipped, or didn't. Last week's issue flagged a release signaled for around September 12th with no , no pricing and no benchmark table. Six days of daily briefs since then show nothing public. Either the launch slipped without announcement or it happened somewhere this desk has not yet looked; both are worth resolving next week.

California's remaining AI bills, due by September 30th. Newsom has signed 13, including Adam's Law's pre-release risk assessments, out of the roughly two dozen bills reported on the governor's desk two weeks ago. The larger measures, including SB 813's third-party safety-assessment framework, remain unsigned with three weeks left.

The Senate's Hugging Face breach deadline, October 1st. Hawley's subcommittee wants answers to 16 questions from Sam Altman by then. GSA's replacement for OpenAI's $1 government pilot, a 27-month deal at half of standard metered pricing, also takes effect October 1st.

Two September 30th prediction deadlines. Anthropic's either flips public by then (Prediction 2026-09-01-F1) or it does not, and OpenAI's ruled-out 2026 listing either holds through the calendar (Prediction 2026-08-06-F2) or Altman's own statement turns out to have been wrong on his own timeline.

The next BLS , on or before October 10th. The first test of whether August's record information-sector job losses were a one-month event or the start of a pattern. Prediction 2026-09-06-B2 settles then.

Sources