Morning Brief, September 30th 2026
Most of Anthropic's compute commitments cannot simply be canceled, Reuters reports. Ted Cruz blocked a vote on mandatory AI safety standards. Europe opened a copyright consultation. EliseAI raised $350 million to expand its housing and health care software.
Anthropic's compute bill comes with obligations it cannot switch off
About 80% of Anthropic's planned $518 billion infrastructure spending is non-cancelable or payable regardless of usage, according to Reuters' further examination of its confidential . The new detail is the contract structure: at least $111.1 billion with Google, $110 billion with Amazon and $31.4 billion with Microsoft, plus roughly $161.2 billion in largely non-cancelable Broadcom equipment leases. The commitments span multiple years.
The contrast is up to $84.5 billion of xAI capacity that is largely cancelable on 90 days' notice. A customer slowdown would therefore reach some suppliers much faster than others. The prospectus remains confidential, preventing independent inspection of these reported terms. SourcesB
Cruz blocks the Senate's shortcut to enforceable AI safety standards
Ted Cruz objected Tuesday when Mark Warner sought unanimous consent to pass a bill creating an AI Safety Board inside the Commerce Department. The proposal would require frontier developers to provide models for review at least 45 days before release, establish safety plans and report incidents. The objection stopped that route to passage; it was not a recorded vote defeating the bill.
Cruz said the proposal gave the board too much power and that he was pursuing a different catastrophic-risk bill. This leaves companies facing proposals for mandatory federal oversight alongside the voluntary agreement announced at the White House. Neither a proposed board nor a promised alternative supplies an enforceable testing regime today. SourcesB
Europe asks whether copyright protection needs another round of measures
The European Commission opened a consultation September 29th on technology's effects on copyright, including the use of protected material in AI. Rights holders, model providers, researchers and consumer groups can respond through November 3rd 2026.
The request seeks evidence and views; it creates no new licensing requirement. It gives creators and developers a concrete opportunity to document where the existing framework fails them. The eventual question is whether further measures follow; the consultation itself does not settle how a particular training dataset may lawfully be used. SourcesA
EliseAI raises $350 million for the software between tenants and landlords
EliseAI raised $350 million at a $4 billion , the company told Commercial Observer Tuesday. It plans a second engineering hub in San Francisco and hiring across engineering, deployment and sales. Its software handles housing and health care operations; the company says it passed $200 million in in June.
The deployment question sits at the other end of the transaction. A tenant asking for maintenance or a patient arranging an appointment needs a completed handoff, even when their request does not fit the usual script. Funding and recurring revenue measure the supplier's position. Neither establishes how often those people get their problem resolved. SourcesB
Efficient Computer raises $97 million to take programmable chips beyond the edge
Efficient Computer announced a $97 million and says its Electron E1 processor is now in volume production. It plans to extend its architecture into applications served by embedded and eventually into data centers.
The technical bet is that energy savings need not require a processor restricted to one workload. Efficient supports familiar programming languages and model formats, allowing software updates as applications change. Its claimed 10 to 100 times energy advantage is the company's own range for general-purpose computation, including AI. It is not an independently established speedup for every model, and the data-center expansion remains a development plan. SourcesA
AutoTrust announces a Qwen-based model for decisions that stay on premises
AutoTrust announced JEV-27B under Apache-2.0 on September 29th. It attaches a trained decision component to a frozen Qwen3.8-27B model, providing structured choices and probabilities while retaining a separate text-generation path. The company says it runs on one Nvidia B200.
This offers a route for organizations that cannot send every routing or classification decision to an outside service. The release includes serving code and evaluation reports. AutoTrust explicitly lists weaknesses in arithmetic, dates, adversarial inputs and multi-step reasoning, and says the model is not intended for high-stakes decisions. Its comparisons with the hosted Jev service are vendor-run evidence. SourcesA
Cloudflare's adaptive AI test finds gaps in its web firewall
Cloudflare described a test in which models proposed variations on known attacks against an authorized staging environment. It recorded 1,107 attempts across six attack categories, then reviewed requests that passed the and used the findings to improve detection.
A request passing the filter was a lead for investigation, not proof that an application had been compromised. The models could neither deploy rules nor change enforcement. Ordinary code restricted destinations, disabled redirects and enforced attempt limits. The useful result is an example of bounded, adaptive testing, with a human verification step between a model's observation and a security finding. SourcesA
Cloudflare uses agents to find cryptography it will need to replace
Cloudflare introduced its internal CryptoLabe workflow for mapping cryptography ahead of its planned 2029 migration. It searches fixed repository snapshots, checks findings against code and traces dependencies that could block an upgrade. Selected prompts are public; the internal application is not.
The company also identifies the missing measurement: it has no dataset for reproducibly comparing prompt versions or claiming complete coverage. Engineers still review the findings. This is a useful deployment account precisely because it separates discovering likely work from proving the inventory is complete. SourcesA
A separate controller helps agents decide where to spend their remaining time
A September 29th introduces an controller that chooses whether to continue a line of work, reuse an earlier result or start again. Workers perform the underlying tasks; the controller maintains a compact account of progress and allocates the remaining budget.
The authors report 71.5% on ProgramBench with GPT-5.5, against 58% for Codex, while gains with another model were smaller. Control overhead can hurt at small budgets. For long tasks, the result makes a practical case for measuring how an agent chooses its next action alongside the quality of the model answering each prompt. These are the authors' evaluations. SourcesA
Agents often abandon the planning structure they announce
Researchers studying the gap between declared plans and execution found that a generic planning-and-action system preserved its intended structure in only 22% to 45% of across three . Routing tasks to executors designed for particular planning patterns improved completion in their experiments.
The remaining failure is choosing the right pattern: the models did not reliably select the best one for each task. A neatly written plan therefore supplies weak evidence about how a run will proceed. The preprint argues for inspecting the actual sequence of actions and enforcing the structure where it matters. SourcesA
A benchmark's simulated customer can quietly make an agent look better
UserProxyBench tests the language model playing the customer in an agent evaluation. Holding the agent fixed and changing only that simulated user shifted mean task reward by 15.2 points across 375 enterprise tasks. Almost a quarter of successful episodes contained a violation of the simulated user's instructions.
The common failure was giving information before being asked. An assistant can then appear efficient because its artificial customer did part of its work. The authors propose scoring user fidelity separately from task completion. Their result challenges comparisons that report an agent score without documenting who supplied the other side of the conversation. SourcesA
KV-Kaizen compresses model memory without immediately discarding the text
KV-Kaizen learns how to compress a language model's working layer by layer. It combines shared caches, lower and reduced representations according to the current context, instead of relying only on removing .
The researchers report a fourfold cache reduction without measured accuracy loss for models of at least seven billion in their tests. A larger reduction on a long-context test also used . The distinction matters: neither result establishes a universal compression ratio. The practical target is keeping a larger model useful within a fixed memory budget. SourcesA
HARISSA tries to keep hard questions local and defer the ones it cannot answer
HARISSA trains a local model to estimate correctness before generating and again after answering. It uses those estimates to decide when extra reasoning is worth the and when to hand the question to a person.
The authors report accuracy within one point of their baseline at 2.7 times lower latency on a single-model device. This offers an alternative to sending every difficult request to a cloud service. The confidence mechanism still makes mistakes: its usefulness depends on the cost of an incorrect answer and whether a person is actually available for deferred work. SourcesA
DexAgent turns a human demonstration into checked robot-training trajectories
DexAgent converts a first-person human video and a task description into simulated robot trajectories. Its stages reconstruct objects, optimize movements and generate training data, using checks tailored to object properties to catch errors before they propagate.
The preprint reports a 3.5-fold success-rate improvement over its comparison systems across eleven real-world tasks. That is evidence for the tested pipeline, not a claim of general robot competence. Its reusable library of tools and checks is the interesting mechanism: a previously solved reconstruction problem can reduce the work needed for the next demonstration. SourcesA
A compiler's access to private code can defeat an agent's read restrictions
A University of Iowa preprint examines a concrete conflict: a compiler must read proprietary modules to build a research program, while the coding agent is supposed to be denied that same source. The author classifies fifteen read routes against a , permission rules and a .
The studied controls do not distinguish which program is reading. That finding does not prove every coding environment leaks, but it identifies a requirement that a broad filesystem boundary alone does not satisfy. Teams protecting source from an assistant need to examine the build path as well as the assistant's direct file-reading tool. SourcesA
World-model verification extends the part of a braking system researchers can analyze
Researchers propose a compact to represent what a vision-based controller sees, then combine several analysis methods to check its behavior. On an emergency-braking benchmark, their procedure resolves the full modeled state space where a prior left 38% unresolved.
On the color-image version, it resolves more than 80%. Resolved means the analysis can reach a conclusion within the modeled setting; it does not mean every state is safe or that a real vehicle has been certified. The advance is reducing the unexamined portion of a perception-and-control system. SourcesA
Character training changes agents' willingness to take resource risks
A new preprint tests whether training around a written can instill risk aversion. The authors report competitive performance against directly trained baselines and better transfer to unfamiliar settings for two of the four models tested.
Their broader hypothesis is that a misaligned agent that fears losing resources might prefer bargaining to a risky confrontation. The experiments support learning a disposition in benchmark decisions; they do not demonstrate prevention of catastrophic behavior. Model choice and training budget strongly influenced the result, limiting any claim that a cautious persona alone is a dependable safeguard. SourcesA
RASO adapts existing agent skills to the software that will actually run them
Retrieval-Augmented Skill Optimization draws from a collection of existing agent instructions, adapts them to a target task and execution system, then revises them using feedback. The September 29th preprint reports gains across four benchmarks and two models against versions without those retrieval stages.
The problem is portability. Instructions written for one set of tools can fail inside another, even when the task sounds identical. RASO treats that mismatch as something to translate explicitly. Its results support evaluating reused instructions after adapting them to the new environment. SourcesA
A builder model gets more value from learned lessons by turning them into tools
Researchers tested a system where one model builds an execution environment for another while both models' stay fixed. A bank of learned principles tells the builder when the target needs support and what resources to provide.
Across two benchmarks, the authors report an 8.95 percentage-point gain over construction without those principles. Giving the same bank directly to the target worked less well. The finding is specific to their evaluation, but it suggests that a lesson can be more useful when translated into executable support than when added to another instruction prompt. SourcesA
Better decisions do not necessarily mean a model has better beliefs
A September 29th preprint separates probabilistic decision errors into two parts: estimating what is true and choosing an action given the costs. On synthetic tasks with known answers, training one part did not reliably improve the other.
Joint training improved both, but depended on matching training and evaluation formats. A buyer evaluating an assistant that makes recommendations therefore needs more than a final decision score. A system can choose better under one cost structure while retaining inaccurate beliefs that become consequential when the task or stakes change. SourcesA
Editorial
An agent's performance score leaves out the person who has to live with its mistakes. The simulated-customer paper gives that problem a measurable form: a cooperative stand-in can volunteer information early and make the assistant seem more capable. The real customer may not know which missing detail matters. They may be calling because the previous handoff already failed. SourcesA
My position is that deployment evaluations should include the cost of recovery to the person receiving the service. For a housing workflow, that means checking whether an unresolved request reaches someone who can act, whether the tenant must repeat it, and whether the software records a success before the work is done. A supplier can measure those outcomes. A procurement team can insist on seeing them.
HARISSA offers one useful direction: defer a question when the system thinks its answer will be wrong. But that only helps when deferral reaches a staffed process. Moving uncertainty into an unattended queue does not make the service safer. Evidence that automated workflows resolve difficult cases without extra effort from the customer would weaken my concern. Until then, the burden of repair belongs in the evaluation. SourcesA
Prediction Watch
No change: Anthropic flips its public by September 30th (Prediction 2026-09-01-F1). The reported document remains confidential; the deadline has not passed. Settles September 30th 2026, at 23:59 UTC. SourcesB
No change: Anthropic's public S-1 confirms the leaked revenue figure (Prediction 2026-09-29-F1). Public verification is still absent. Settles March 31st 2027. SourcesB
Nothing settled during this reporting run. The two September 30th calls remain open until their deadlines can be checked. No fresh Chinese flagship release was verified; AutoTrust's Qwen-derived decision model is covered above.
Sources
- B Anthropic's compute bill comes with obligations it cannot switch off
- B Cruz blocks the Senate's shortcut to enforceable AI safety standards
- A Europe asks whether copyright protection needs another round of measures
- B EliseAI raises $350 million for the software between tenants and landlords
- A Efficient Computer raises $97 million to take programmable chips beyond the edge
- A AutoTrust announces a Qwen-based model for decisions that stay on premises
- A Cloudflare's adaptive AI test finds gaps in its web firewall
- A Cloudflare uses agents to find cryptography it will need to replace
- A A separate controller helps agents decide where to spend their remaining time
- A Agents often abandon the planning structure they announce
- A A benchmark's simulated customer can quietly make an agent look better
- A KV-Kaizen compresses model memory without immediately discarding the text
- A HARISSA tries to keep hard questions local and defer the ones it cannot answer
- A DexAgent turns a human demonstration into checked robot-training trajectories
- A A compiler's access to private code can defeat an agent's read restrictions
- A World-model verification extends the part of a braking system researchers can analyze
- A Character training changes agents' willingness to take resource risks
- A RASO adapts existing agent skills to the software that will actually run them
- A A builder model gets more value from learned lessons by turning them into tools
- A Better decisions do not necessarily mean a model has better beliefs