Morning Brief, September 18th 2026
Anthropic says Claude now leads a quarter of its model research. The company confirmed it operates a biology lab. Figure tested household robots in unfamiliar homes. Singapore announced Waymo’s route toward commercial rides in 2028.
Anthropic says Claude leads a quarter of its model research
Anthropic reports that Claude led 26% of its AI research and development work in August, meaning it could carry out most of a task from a high-level instruction under human supervision. No measured category was fully autonomous. The company built the index using internal work records and model-assigned ratings, with human comparisons.
The disclosure supplies a repeatable measure of automation inside a frontier lab. It does not measure the acceleration of scientific progress directly: completing more of an existing task basket and discovering better research directions are different achievements. The company also reports monitoring and computing-allocation measures, while acknowledging that comparable reporting across labs needs a common methodology and independent checks. SourcesA
Anthropic confirms it operates a physical biology lab
Anthropic’s life-sciences head Eric Kauderer-Abrams confirmed to Reuters that the company conducts biology experiments in its own facilities as well as through outside partners. Reuters places the in the San Francisco Bay Area. A spokesperson clarified that the facility is not specifically a drug-discovery lab.
Bringing experiments inside the company shortens the route from a model’s proposed experiment to a physical result. It also complicates relationships with pharmaceutical customers whose research the company helps power. Anthropic says it walls off customer data and is not running clinical trials. The report establishes a laboratory, not a treatment ready for patients. SourcesB
Verified research teams gain broader access to Claude’s biology capabilities
Anthropic opened applications for its Life Sciences Verification Program after enrolling dozens of organizations in early access. The beta gives approved teams more permissive biology safeguards on Mythos, Opus and Sonnet. A separate project-specific permission tier covers higher-risk work; broader access to that tier for Mythos remains restricted.
Applicants must establish research , security practices and ethical oversight. The program requires 30-day data retention for monitoring; other safeguards remain in place. The opening changes who can attempt previously blocked scientific work, but enrollment totals alone do not show which smaller institutions can obtain and use access. SourcesA
Figure reports 56% task success in homes its robot had never seen
Figure tested Helix 2.5 on room tidying, towel folding and bed making across 30 unfamiliar Bay Area homes. In its comparison, on the company’s Index human-behavior dataset raised complete-task success from 9% to 56%, with other experimental conditions held fixed. Evaluation homes supplied no adaptation data.
The improvement addresses an expensive deployment requirement: collecting fresh training data wherever a robot will work. But nearly half the tested tasks still failed. These are Figure’s measurements across selected behaviors, not evidence that a general household robot can work unattended. SourcesA
Singapore confirms Waymo’s route toward a 2028 commercial launch
Singapore’s Land Transport Authority says Waymo vehicles should arrive in the coming months, begin initial manual driving in 2027 and aim for commercial ride-hailing in 2028. Waymo joins an autonomous-vehicle landscape that already includes ComfortDelGro and Grab.
The authority says deployments must meet safety and regulatory requirements and will not outpace support and retraining for affected drivers. That makes workforce transition an explicit part of the rollout. The announcement is a development schedule; riders cannot yet summon this service. SourcesA
Claude’s biology software optimizations come with downloadable code
Anthropic released optimizations for more than 30 biomolecular models, reporting roughly fourfold average speedups with a small tradeoff. It also describes a lower-memory mode for larger structures. The distinction between running a calculation and obtaining a correct structure matters: its largest demonstration systems ran, but their predicted structures were inaccurate.
The company is also backing experimental validation through a protein-design competition with Adaptyv Bio. Faster computation lowers the cost of testing candidate designs; laboratory results must establish whether the candidates work. The code is a reference release without planned maintenance. SourcesAA
Canada and Germany plan public funding for LawZero’s alternative to autonomous agents
Canada and Germany announced plans to invest CAD 150 million and EUR 100 million respectively in LawZero. The September 16th announcement supports its Scientist AI approach, which aims to provide evidence-based answers without pursuing goals of its own. Germany’s funding remains subject to notification to the European Commission.
This funds an architectural alternative alongside the commercial race to build more autonomous systems. The program initially targets oversight tools and scientific research. Governments’ stated safety ambitions are not a demonstrated safety guarantee; that requires testing the resulting systems. SourcesA
Mantic raises $25 million after AI forecasters lead a tournament
Mantic announced a $25 million seed round led by Radical Ventures, with participants including Microsoft’s M12 and Thinking Machines Lab, Reuters reports. The London company specializes other labs’ for forecasting. Reuters says its probabilities outperformed human competitors in the summer Metaculus Cup.
A forecasting tournament supplies something an eloquent prediction cannot: outcomes against which probabilities can be scored. The commercial test is transfer to customers’ decisions, including questions with thin evidence or incentives to manipulate the forecast. The funding was not disclosed. SourcesB
Chinese AI shares face pressure over access to US models
Z.ai and MiniMax shares had each fallen more than 30% during September, Bloomberg reported on September 17th, cutting their combined market value by $33 billion. Its report identifies competition and cash consumption alongside concern about US developers’ calls to restrict Chinese access to their models.
The disputed practice is : using another model’s responses as training material. A stock decline does not establish either misconduct or an enacted restriction. The business exposure is to access terms and policy decisions that could raise the cost of improving competing models. SourcesB
Tower and NewPhotonics move integrated optical engines into volume shipments
Tower Semiconductor and NewPhotonics announced volume shipments of optical-engine chips with integrated lasers, supporting connections at 800 gigabits and 1.6 terabits per second. Their faster 6.4-terabit products remain scheduled for volume shipments in the first half of 2027.
The manufacturing step matters because large AI systems must move data between processors as well as compute on it. Integrating optical components can simplify assembly, although the announcement supplies no independently measured system-level energy savings. The current shipment and the faster roadmap product should stay separate in buyers’ capacity plans. SourcesA
Amazon prices AI interviews by completed candidate evaluation
Amazon Connect Talent lists a $20 charge per completed candidate evaluation, combining an AI-led interview with an assessment. AWS says there are no recruiter-seat fees or minimum monthly commitments, and abandoned evaluations are not billed.
The price gives hiring teams a direct comparison against screening labor and existing software. It does not establish that the scores predict job performance or treat applicants fairly. An interview transcript and audit trail let a recruiter inspect what happened; they do not by themselves validate the hiring decision. SourcesA
DeepMind opens an institute for arguments about the world after AGI
Google DeepMind’s new institute publishes essays on reasoning transparency, economic policy and frontier-AI governance. Its opening collection includes work by Demis Hassabis, Shane Legg and James Manyika, alongside other researchers.
The institute explicitly says these essays express their authors’ ideas and should not be treated as Google’s official position. That boundary matters when a paper proposes an oversight mechanism: publication starts an argument, while a corporate commitment would require an accountable decision and implementation. SourcesA
Tanium tells customers to expect more patches as AI finds more flaws
Tanium says its work with frontier models, including Claude Mythos 5 through Project Glasswing, will increase the frequency of and the number of disclosed vulnerabilities. Its existing monthly maintenance-release commitment remains in place.
The operational consequence falls on customers’ patching teams. Discovery can accelerate faster than deployment of fixes, leaving organizations with more known exposure even as the vendor improves detection. Tanium’s announcement makes that workload visible; a larger count alone cannot determine whether a product became less secure. SourcesA
HP finds fake AI trading agents replacing crypto-wallet extensions
HP’s latest threat report describes attackers using supposed AI trading to persuade users to download malware. The software replaces trusted cryptocurrency browser extensions with malicious versions.
The AI connection here is the lure. The report does not establish that an autonomous model performed the theft. That distinction directs the defense toward software and browser-extension integrity instead of treating every incident marketed with an AI label as a new model capability. SourcesA
ScienceIDE turns scientific repositories into agent-training environments
ScienceIDE converts scientific code repositories into executable tasks with expert-defined acceptance criteria. Its authors use verified interaction records to train a family of models and report gains on scientific-code repair as well as selected general . Code is public.
The contribution is a way to make specialized software teachable and testable. A program that runs is not necessarily a valid scientific analysis, so the value depends on the domain-specific checks. The reported gains remain the authors’ evaluation, with generalization beyond those tasks still to be established. SourcesA
Robots learn from demonstration videos without a fresh training run
GPT-Policy combines a with a controller that checks and executes proposed robot actions. In real-robot trials, its authors found that human demonstration videos improved task completion even without matching robot-action labels. Aligned action references supplied further help on tasks involving contact.
The proposed route to adaptation is context supplied at deployment, without changing the model’s learned . That could reduce retraining work for each new task. The paper evaluates tested scenarios and their limitations; it does not establish safe execution of arbitrary instructions. SourcesA
A materials agent tests experiments intended to disprove its own explanation
SynAgent’s materials-synthesis researchers report a campaign of 18 autonomous experiments in which the system revised explicit hypotheses about growth. It selected conditions expected to fail as well as conditions expected to succeed, using measurements to challenge its explanation.
The output includes a testable account of the process, not just an optimized sample. That gives a scientist something to check and reuse. Evidence from a single material and campaign cannot establish a general autonomous scientist, but the deliberately disconfirming experiments make the result more informative than a best-sample demonstration. SourcesA
Physical world models need separate fixes for drift and changed conditions
A new study separates two failures in learned physical simulators: predictions drifting over long sequences, and predictions failing when a physical relationship changes. The authors show that respecting energy-conserving motion improves stability, while explicitly representing how components interact helps the model adapt when that interaction changes.
Removing either component damaged its corresponding capability without necessarily destroying the other. That controlled separation gives builders a more precise diagnosis than a single prediction score: a stable simulation can still answer a changed-world question incorrectly. SourcesA
Agent recovery improves when rollback preserves lessons from the failed attempt
Rollback-Induced Reflection restores an agent’s environment to an earlier state while keeping selected lessons from the abandoned attempt. Its authors report improvements across three long-task benchmarks and multiple underlying models.
The design addresses a practical mismatch: rewriting an agent’s context does not undo changes it already made, while restoring everything can erase the knowledge needed to avoid the same mistake. Its usefulness depends on an environment that can actually be restored. A sent message or an irreversible physical action cannot be recovered by editing memory. SourcesA
Editorial
The cost of useful automation includes the work left after the model finishes. A cheaper candidate interview still needs a defensible hiring decision. A robot that completes a household task only some of the time still needs someone available for the failures. Buying the successful demonstration without budgeting for those obligations is how a capability gain becomes an operating loss.
My position is that buyers should compare total cost per accepted outcome, including review, recovery and the staff required when automation stops. The vendor’s unit price is a starting point. Procurement should require a trial in which exceptions count against the result and the customer records the labor spent resolving them.
There is a competitive consequence. A supplier that publishes failure conditions helps a customer calculate a budget. A supplier that publishes only its best outcome forces the customer to discover the liabilities after purchase. I would pay more for the former when the task carries a costly failure. That preference would be wrong if broad customer trials showed that recovery costs remain negligible even as systems encounter unfamiliar conditions. Today’s household-robot result makes that a claim to test before committing to unattended operation. SourcesAA
Prediction Watch
No change: none of the capability gates publish small-institution numbers (Prediction 2026-09-06-T1). Anthropic reports dozens of early-access organizations, without the enrolled-and-active breakdown for the specific kinds of small institutions named in the call. Settles March 31st 2027. SourcesA
No change: Anthropic’s embedded evaluators publish no incident report in year one (Prediction 2026-09-13-T1). The new internal measurements come from Anthropic itself, rather than an evaluator operating under its embedded-access program. Settles September 12th 2027. SourcesA
Nothing settled in the evidence reviewed today. The China/ sweep found financial pressure on Chinese labs, but no verified new downloadable release meeting the large-model call’s threshold.
Sources
- Anthropic’s research automation measurements
- Anthropic’s biology lab, Reuters
- Life Sciences Verification Program
- Figure’s Helix 2.5 evaluation
- Singapore’s transport authority on Waymo
- Anthropic’s biomolecular modeling optimizations
- Mantic’s seed round, Reuters
- Canada and Germany’s planned LawZero investment
- Chinese AI stocks and access restrictions, Bloomberg
- Tower Semiconductor and NewPhotonics shipment announcement
- Amazon Connect Talent pricing
- DeepMind Institute
- HP’s threat research on fake trading agents
- Tanium’s frontier AI security commitment
- RebarSim paper
- ScienceIDE paper
- In-Context Robot Learning with VLM Agents
- Hypothesis-Driven Autonomous Materials Synthesis
- Stability and counterfactuals in physical world models
- Rollback-Induced Reflection paper
- Biomolecular optimization repository and maintenance status