Morning Brief, September 13th 2026
Sam Altman ruled out an OpenAI IPO this year. Microsoft reportedly plans to triple its data-center capacity. Massachusetts tied large data-center permits to community agreements. Nvidia released the training recipe behind its Olympiad proof system.
Altman rules out a 2026 OpenAI IPO as safety work takes priority
Sam Altman said OpenAI will not go public this year, citing the company's safety work, in a Fortune interview published Saturday. Bloomberg reports that he put a listing in next year rather than the remaining months of 2026.
That removes the September offering from management's stated plans. It does not supply a binding 2027 date, a public or evidence that the underlying safety problems have been resolved. Investors waiting for public financial disclosures still have to wait, while customers have a separate question: which development and deployment decisions will actually change? SourcesB
Microsoft reportedly plans to triple its data-center capacity by 2032
Microsoft plans to expand its data-center capacity from about 12 gigawatts to more than 38 gigawatts by 2032, Bloomberg reported September 10th, citing people familiar with the plans. Its updated report says the total includes owned and leased facilities and excludes capacity rented from .
The figures describe electrical capacity across the estate, not a threefold increase in AI performance. They nevertheless put a scale on Microsoft's response to shortages that have forced it to turn away some cloud and AI business. A plan this large depends on power connections and construction schedules as well as chip deliveries. SourcesB
Massachusetts makes large data-center permits conditional on community benefits
Massachusetts Governor Maura Healey's September 8th executive order directs state agencies to withhold permits for data centers exceeding 25 megawatts of peak demand unless applicants meet the state's development framework and submit a . It also directs regulators to establish a payment mechanism by December 31st for facilities that do not procure enough additional clean electricity.
This gives permitting agencies a condition they can apply before construction. The fee mechanism still needs implementation. Developers evaluating sites must account for the community agreement and additional electricity requirement alongside the cost of land. SourcesA
Adobe reports $6.76 billion in quarterly revenue as AI subscriptions grow
Adobe reported fiscal third-quarter revenue of $6.76 billion on September 10th, up 13% from a year earlier, or 12% with exchange-rate movements removed. Its earnings release says annualized recurring revenue from products it calls AI-first grew more than 150%; total company reached $27.50 billion.
The release supplies evidence of subscription growth at an incumbent selling AI tools. It does not isolate the profit from those tools or establish how much growth came from customers paying more versus customers buying for the first time. The total ARR and AI-first growth rate measure different things and cannot be substituted for one another. SourcesA
FANUC schedules a drawing-to-welding agent for December shipments
FANUC announced an AI Welding Agent on September 11th that uses Google Cloud technology to read component drawings and generate robotic arc-welding programs. It plans a demonstration at the International Welding Show beginning September 16th and shipments at the end of December.
The system targets the preparation work between receiving a drawing and teaching a robot its motions. That is a specific industrial job with a measurable result: whether the resulting weld meets the specification. The announcement establishes a product schedule; it does not yet establish unattended reliability on customers' production lines. SourcesA
SenseNova publishes the training account behind its downloadable image model
SenseNova released the U1.5 technical report this week, explaining how its model combines visual understanding and image creation. The project's release history dates the base to August 20th and the report announcement to September 11th. This is new documentation of an existing .
The weights carry an license. The paper describes combining specialized models for editing, aesthetics and bilingual text rendering, and promises training-code release. Its maintainers still list dense-text errors and drift during complex edits among known limitations. Builders can inspect and run the model now without treating the promised training pipeline as already delivered. SourcesAAA
Nvidia releases the checkpoints and data behind its Olympiad proof pipeline
Nvidia researchers' September 9th paper describes a Nemotron system that scored 30 of 42 points on the 2026 International Mathematical Olympiad, meeting the gold-medal threshold. They release specialist checkpoints, training data, code and submitted solutions, plus a new collection of Olympiad-level problems.
The pipeline generates, checks and refines proofs in natural language. It uses no formal prover or internet access. The useful development is the reproducible training and inference recipe; a natural-language verification pass still does not provide a machine-checked certificate that the proof is valid. SourcesA
Pocket Entertainment reports a $500 million revenue run-rate
Pocket Entertainment, the parent of Pocket FM and Pocket Saga, announced September 10th that its revenue exceeded $500 million, with 70% year-over-year growth. It attributes part of the expansion to AI used across production, localization and distribution.
The company says AI has reduced the time required to enter new markets from approximately twelve months to two. These are company measurements, and the run-rate is a projection of current revenue, not revenue recognized over a completed year. The commercial test is whether faster adaptation of stories can sustain paying audiences once the initial catalog expansion slows. SourcesA
Instacart brings conversational carts to customers and grocers' own sites
Instacart launched Clementine on September 9th for most US and Canadian customers. The assistant turns requests and recipes into carts. In the same announcement, it named Food Bazaar, Heritage Grocers Group and Woodman's as live users of Cart Assistant, its retailer-branded version.
The retailer product connects the assistant to the grocer's catalog and customer data. That lets stores offer the shopping interface on their own sites and apps. A generated cart still needs customer review, and availability is not universal: Instacart's help page explicitly says Clementine is not available to everyone. SourcesAA
EvoSafeHarness tailors agent defenses to the model and its job
EvoSafeHarness, a September 5th preprint, searches for a combination of written policy and executable controls suited to a fixed model in a particular domain. It tests candidate defenses against attacks while measuring how much useful work they prevent.
On DecodingTrust-Agent, the authors report reducing average attack success from 45.6% to 10.0%, at a 3.3-point utility cost.
The result supports testing defenses against the exact model and application that will use them. It does not demonstrate that has been solved, especially outside the evaluated tasks and attack budgets. SourcesA
IdeaAMBIG finds that models miss the instructions a researcher left out
IdeaAMBIG, released September 9th, tests whether a research specification contains enough information to implement the intended method. Its cases combine gaps found in reproduction reports and GitHub issues with controlled omissions from complete specifications.
Across the tested models, the best score for finding defects in real cases was 9.6%. When supplied the defect, the best model's clarification-action score was 80.6%. The measures test different tasks, but their separation identifies a practical failure: an agent can ask a useful question after someone tells it what it missed. Delegating implementation still requires someone to check that the specification says enough. SourcesA
A solver can approve the wrong translation of a problem
A September 10th paper identifies a failure in automated : an incorrect translation can execute successfully and return the expected verdict while failing to represent the intended problem. A check on the solver's final answer alone can therefore miss the error.
The authors train a verifier to estimate whether a translation matches a reference formalization, using examples checked with the solver. They report improved detection and downstream accuracy. This is an experimental safeguard for the translation step. It leaves the solver's guarantees intact while showing why those guarantees do not automatically extend to the sentence supplied by a user. SourcesA
Cohere researchers release a small model that reasons in the user's language
Cohere Labs researchers' September 9th paper introduces Tiny Aya L2-Thinker, a 3.35-billion-parameter model designed to keep its reasoning in the language of the prompt. They report an in-language reasoning rate above 93% across 60 languages and release weights and multilingual reasoning data.
Their method combines broader language coverage with English reasoning examples and multilingual data that does not itself contain reasoning. That reduces the need to obtain worked reasoning examples in every language. Staying in the requested language is an accessibility result; it is not a claim of 93% answer accuracy. SourcesA
A transit-kiosk benchmark gives small local models a concrete job
MetroLLM-Bench, published September 9th, tests routing, fare calculation, accessibility and other kiosk decisions across real transit systems. Models must call structured tools and return a state that the kiosk can render, including a fare quote when appropriate.
A tuned Qwen 3.5 model with a 2.6 GB footprint scored 91.3 on the portion of the test, against 84.6 for a rule-based baseline. This suggests a narrow task where local language models deserve comparison with conventional software. It does not prove that a working kiosk can recover from every payment failure or stale timetable; operators must test those conditions separately. SourcesA
XPeng's speech-model compression preserves most accuracy with fewer layers
XPeng researchers' September 10th X-AuT paper describes removing layers from a speech model's in stages, then restoring performance using a larger teacher. Removing layers all at once can cause deleted words and premature endings.
Their fourteen-layer version has 20.7% fewer audio-encoder than the eighteen-layer starting point, with mean error of 5.75% versus 5.61% across the Chinese-English tests. The authors label the results single-run measurements. This is a possible reduction in the audio component's cost, not a measured saving of the same size across the entire speech service. SourcesA
Negative Self-Distillation trains models away from their own bad reasoning
A September 10th preprint proposes generating flawed reasoning with a model and training it to move away from those behaviors. The authors argue that imitating a solution produced with advance knowledge of the answer can suppress the uncertainty and self-correction needed for difficult problems.
Their method selectively targets reasoning-related tokens because indiscriminately penalizing the flawed response can also damage ordinary language ability. They report gains over the evaluated self-training baselines without external answer labels. The result is evidence about a particular training method, not proof that a model can reliably judge all of its own mistakes. SourcesA
Mi-Ripple tackles the texture damage left by repeated AI image edits
The September 10th Mi-Ripple preprint describes a restoration workflow for grid-like and granular artifacts that accumulate during repeated AI image editing. It separates regular patterns that can be filtered from textures entangled with legitimate detail, then uses cleaned references when regeneration is necessary.
The evidence is small: the paper reports fourteen filtering executions and a paired regeneration example. It is an early workflow to inspect if successive edits are degrading an image, with no broad claim that restoration preserves every original detail. Regeneration can replace information, so a cleaner picture still needs comparison with the source. SourcesA
Editorial
Conviction: a useful AI test must check whether the system solved the intended problem. That sounds easy until the work passes through several translations. A person describes a task. A model supplies code or a formal statement. A tool returns success. Each step can be internally consistent while the first instruction gets lost.
IdeaAMBIG makes the first failure concrete: models were much better at producing a clarification after being shown a defect than at locating the defect themselves. The solver-verification paper catches the next failure: a wrong translation can still produce the expected verdict. Adding a checker helps only if it checks the property the user actually needs. SourcesAA
The same question belongs in a factory. FANUC's proposed welding agent moves from a drawing to robot motion. A program that runs without an error is one checkpoint. A weld that meets the drawing is another. Buyers should require a route from the original specification to the final inspection before paying for unattended operation. SourcesA
Speculation: vendors that make this chain inspectable will win more demanding deployments even when their underlying model is weaker. A buyer can improve a measured failure. A plausible success with the wrong objective is harder to detect. The evidence that would overturn this view is sustained production use in which general-purpose agents meet independently inspected specifications without that additional checking process.
Prediction Watch
More likely now: OpenAI does not IPO in September. Altman's public statement that no listing will happen this year supports the call, but the criterion concerns actual trading through the deadline. The call stays open (Prediction 2026-08-06-F2). Settles September 30th 2026. SourcesB
No change: Nobody solves prompt injection. EvoSafeHarness reports better defenses on evaluated tasks with a cost to useful work. It does not establish a broadly adopted architectural solution (Prediction 2026-08-06-T4). Settles August 6th 2027. SourcesA
Nothing settled today. The checked sources establish neither a completed OpenAI listing nor a generally adopted fix for prompt injection.
Sources
- B OpenAI IPO timing, Bloomberg
- B Microsoft capacity plans, Bloomberg
- A Massachusetts Executive Order 658
- A Adobe fiscal Q3 earnings release filed with the SEC
- A d-Matrix rack integration announcement
- A FANUC AI Welding Agent announcement
- A SenseNova release history
- A SenseNova U1.5 technical report
- A SenseNova U1.5 model card
- A An Open Recipe for IMO Gold
- A NCP-ArchPreview technical report
- A Pocket Entertainment growth announcement
- A Instacart launch and retailer rollout
- A Clementine availability and use
- A Cymphony launch release
- A Cymphony CEO account
- A EvoSafeHarness paper
- A IdeaAMBIG paper
- A Beyond Solver Verdicts
- A Building Multilingual Bridges
- A MetroLLM-Bench paper
- A X-AuT paper
- A Negative Self-Distillation paper
- A Mi-Ripple paper