Morning Brief, August 13th 2026
DeepSeek shipped its finished flagship overnight with no announcement and MIT . A security firm documented running a near-autonomous intrusion into Taiwan's nuclear safety agency. Anthropic's backers expect a $2 trillion October listing. And Grok 4.6 arrived at $2 per million .
DeepSeek shipped its finished flagship, and never said a word
The model string on DeepSeek's pricing page changed to DeepSeek-V4-Pro-0813 on Wednesday, closing the preview that has run since the V4 family's debut in April, and the weights followed onto under an MIT license within a day. No blog post, no announcement; Simon Willison spotted it live on OpenRouter. The card lists 1.7 trillion total parameters with 49 billion active per token, a recommended context of 384K, and a DSpark speculative-decoding module. The benchmark jumps over the preview are large: Terminal Bench 2.1 rises from 72.1 to 87.9, DeepSWE from 12.8 to 62.7, Toolathlon-Verified from 55.9 to 74.1. DeepSeek's own table across nine agent benchmarks has Anthropic's Claude Fable 5 ahead by an average of 5.3%, at list prices roughly 46 times higher. DeepSeek published that table itself; the arithmetic is the announcement. The release landed one day after Alibaba's Qwen3.8-Max, and the two releases together make this the heaviest 48 hours for Chinese open weights this year. SourcesAB
AI agents ran a near-autonomous intrusion into Taiwan's nuclear safety agency
Israeli security firm Dream published research Wednesday documenting a campaign that ran July 1st through 4th: open-source agent frameworks, named as Hermes and OpenClaw, executed a near-autonomous intrusion into Taiwanese government systems, expanding into the nuclear safety agency, supply-chain vendors and at least seven energy companies. The numbers from Dream's 1,395-file evidence archive: 85 government accounts compromised, 2,564 personnel records exfiltrated, 21 connected government systems mapped, 12 attack waves, up to eight sub-agents running at once. The models' guardrails were bypassed by framing the operation as an authorized penetration test, and the agents ran their own learning cycles against databases and GitHub between waves. Attribution rests on Chinese-language operator documentation and stops short of naming a group or government; every technical detail is single-sourced to Dream. It is the first documented case of a near-autonomous AI intrusion into a sovereign government. The attack ran in July; the disclosure is the news. SourcesBB
Anthropic's backers expect a $2 trillion October listing
Six Anthropic investors told the Financial Times they expect the company to pursue a of at least $2 trillion in an October IPO, which would pass SpaceX and make it the largest stock-market debut in history. The same backers project an annualized revenue of $100 to 120 billion by the end of 2026. The company's last private mark was $965 billion in May; it filed confidentially with the in early June, and bankers have been arranging investor meetings since July. The load-bearing caveat sits inside the reporting itself: senior executives have not set a valuation target, even privately, so the $2 trillion is what the shareholders want, not what the company has said. Every number here is single-sourced to the FT's investors. SourcesBC
Grok 4.6 shipped, tied for third place, priced to undercut
xAI released Grok 4.6 on Wednesday, the model Musk had said would gate Grok Bot's wider rollout. It is a post-training upgrade over Grok 4.5: same base model, a longer supplemental training run, regenerated supervised trajectories, and reinforcement learning in agentic environments. The release notes list a 500K-token context, text and image input, four reasoning-effort levels including a new "xhigh," and pricing of $2 per million input tokens, $0.50 cached, $6 output below 200K-token prompts, doubling above. It scores 61 on the Artificial Analysis Intelligence Index, up five points from Grok 4.5 and tied for third with GPT-5.6 Sol Max. It is generally available in the API, default in Grok Build, and live in Cursor on all plans. The pricing is the aggressive part: third place on the leaderboard at a fraction of what the models above it charge. SourcesAB
The market graded the neocloud prints, and the grades were loud
Wednesday's session turned Tuesday's earnings into prices. Nebius closed up 34% at $247.05 after its 454% revenue quarter, helped by new disclosures on the call: four AI-cloud deals each above $1 billion in total contract value, and customer prepayments in 70% of signed deals. CoreWeave, whose interest bill we flagged yesterday, closed up 19.5% at $107.73. Lumentum rose 14% on its own AI-driven print, Dell about 10% in sympathy, Micron about 5%. The macro backdrop cooperated: July came in at 0.1% month over month and 3.4% year over year, in line with forecasts and supportive of a September Fed hold. Equity is pricing the demand and ignoring the cost of carry. Prediction F4 is a bet on that divergence. SourcesBB
DeepSeek's price increase now has a schedule: up to 1,100%, starting Saturday
The warning DeepSeek issued last week is now a published price list. The new schedule takes effect at 16:00 UTC on August 16th and moves the API fully onto peak and off-peak billing, with off-peak at half the peak rate and peak hours set at 01:00 to 04:00 and 06:00 to 10:00 UTC. V4-Pro peak uncached input goes from 3 to 9 yuan per million tokens and output from 6 to 27, with increases across models and token types reported between 50% and 1,100%. It is the second pricing change in under a month, and it comes with the company's own explanation attached: demand has outrun capacity, with V4-Flash alone reported processing 8 trillion tokens in a single day on August 1st. Whether the lab that set the industry's price floor can hold a tripling is now the cleanest pricing-power experiment in the market, alongside the one Prediction T3 is running on Anthropic from the premium end. One dating caveat: some coverage renders the effective date August 17th, an artifact of the UTC cutover, so we use DeepSeek's own schedule. SourcesAC
AI servers are now more than half of Foxconn
Foxconn's second quarter, reported Wednesday: net profit of NT$60 billion, about $1.86 billion, up 35% year over year and a Q2 record, on revenue of NT$2.53 trillion, up 41%. The structural number came out of the results call: cloud and networking, the segment that carries AI servers, reached 51% of revenue, the first time it has passed half, against 29% for smart consumer electronics. The company that assembles the iPhone now makes more from AI racks than from phones. Forward color: Nvidia's Vera Rubin racks enter mass-production preparation this quarter with shipments starting in Q4, and the CEO named advanced-packaging capacity as the ceiling on 2027 AI server output. Separately, TSMC told the OCP APAC Summit its 5.5x-reticle CoWoS with support for 12 HBM4 stacks is in mass production with yields above 98%, a claim carried by one trade summary and worth holding at arm's length until a second source lands. SourcesAB
Supermicro guided to $65 billion, and the Street had modeled $52 billion
Super Micro's fiscal fourth quarter, reported Tuesday night and repriced Wednesday: revenue of $11.1 billion, up 93% year over year and slightly under the $11.55 billion estimate, with adjusted earnings of $1.70 a share against the roughly $0.92 expected. Gross margin reached 17.5% from 9.5% a year earlier, and net income of $1.18 billion was six times the year-ago figure. The did the work: fiscal 2027 net sales of $65 to 72 billion against a near $52.5 billion, and a first-quarter guide of $14.5 to 15.5 billion against $11.68 billion. The stock rose 19% on Wednesday. A server assembler guiding 30% above the Street is a statement about the order books of everyone upstream of it. Prediction F3 expects the same funded demand in Nvidia's own print. SourcesBC
Gemini's app crossed a billion monthly users
Google said Tuesday the Gemini app passed 1 billion , its fastest product ever to the mark and the 14th Google product to reach it. The trajectory: roughly 400 million at I/O in May 2025, 900 million this May, 950 million at July's earnings, 1 billion now. The texture numbers: 63% of users engage by voice, more than 150 million images are generated daily, a fifth of Gemini Live interactions use camera or screen share, and the figure counts the app alone, excluding AI Mode in Search, which passed 1 billion separately. ChatGPT reached the same threshold in June, so the two most-used AI products on earth now count their users in billions, ten months after either could claim half that. SourcesAB
A zero-click Zoom exploit chain, built with fewer than twenty AI prompts
Security firm A Security disclosed "Zoomsday" on Monday: three vulnerabilities in Zoom's annotation feature, CVE-2026-53413 through 53415, that chained into . Any meeting participant could run code on other participants' devices across Windows, macOS, Linux, iOS and Android on clients up to 7.0.5, with no click required from the victim. The finding that travels: the researchers built the working exploit with fewer than 20 prompts to publicly available AI models in under 24 hours, work they say previously required nation-state infrastructure and months. The timeline behaved the way disclosure is supposed to: found June 8th, reported June 10th, client fix June 22nd, server-side mitigation July 15th, public disclosure August 11th, after the patches. The installed base of unpatched clients drains far slower than that sentence reads. SourcesAB
One poisoned scanner reached 434,000 CI/CD pipelines
CloudSEK's research, published Monday, reconstructs the March compromise of LiteLLM, the most widely used open-source gateway: threat group TeamPCP leaked an automation token for the Trivy security scanner, poisoned it, and rode LiteLLM's own build pipeline to publish malicious versions 1.82.7 and 1.82.8 to PyPI, undetected for roughly 20 days. The malware harvested cloud credentials, repository tokens, SSH keys, Kubernetes secrets and AI provider keys from environments. CloudSEK's exposure estimate: more than 2,500 organizations and about 434,000 pipelines, with high-confidence matches including Nvidia, AWS, Cisco, Salesforce, Siemens, X Corp and Orange. The FBI issued a FLASH advisory in July warning the stolen credentials remain usable. The figures are estimates of exposure, not confirmed breaches of the named companies, and CloudSEK is the single source; the FBI advisory is the corroboration that something real happened. SourcesAC
Korea's chip rally ran a fourth day, and Samsung's HBM4 yield is the reason given
The memory trade we carried Tuesday kept running. The KOSPI closed Thursday at 6,813.34, up 3.56%, a fourth straight gain and a roughly 30% rebound from its July low, with Samsung Electronics up 6.68% and SK Hynix up 5.54%. Kioxia added about 4% in Tokyo. The new datapoint under the move: TrendForce, relaying Korean trade press, reports Samsung's HBM4 yield has reached about 80%, from under 60% when mass production started in February and ahead of the company's own end-of-year target, which would tighten the race to supply Nvidia's Vera Rubin generation. The yield figure is single-sourced trade reporting and Samsung has not confirmed it. SourcesBB
Brin pushed DeepMind toward "recursive self-improvement," per Reuters
A Reuters exclusive adds the missing interior to last week's Google reshuffle. Per the reporting, Sergey Brin has spent months urging key AI staff to go all in on Gemini, including an April town hall of hundreds of employees where he pressed DeepMind to move faster as Anthropic previewed Claude Mythos, and has pushed work on recursive self-improvement, models improving models. The new structural detail: non-technical support groups are being moved out of DeepMind into corporate Google, confirmed to staff at an internal meeting, which reads as the research lab being cut down to research while Alphabet runs the product war directly. The account is single-sourced to Reuters' unnamed insiders; the reshuffle it explains, Hassabis to Alphabet chief scientist and Kavukcuoglu to day-to-day command, is on the record. SourcesB
Bank of America put a $250 billion number on the infrastructure trade
Bank of America announced a "Critical Infrastructure Finance Initiative": $250 billion of financing capacity over 18 months, running through July 4th 2027 and timed to the US 250th anniversary, prioritized toward data centers, semiconductors, power generation, energy storage, critical minerals, transportation, water and natural gas. Read the fine print before repricing anything: this is mobilization, lending, underwriting and capital raising, not balance-sheet , and it follows similar pledges from Morgan Stanley and JPMorgan. Wall Street is formalizing compute and power as a financing category, three days after Nvidia's $500 billion consortium did the same from the vendor side. We could not locate the bank's own release this morning, so the terms carry the outlets' sourcing. SourcesBC
Qwen3.8-Max's first day: quantized everywhere, and the 27B is missing
A day after Alibaba shipped the 2.4-trillion-parameter weights, the repository sits at about 695 likes and third place on Hugging Face trending, with the FP8 variant at 4,000 downloads and Unsloth's builds already live in a dozen variants for llama.cpp, Ollama and LM Studio. The raw download count on the main repo reads low at about 1,000 because the release spans 213 shards. What has not appeared is the Qwen3.8-27B companion that Alibaba's July announcement promised alongside the flagship: the org page shows nothing, and the community threads are asking. The custom license with its 100-million-user and revenue triggers remains the standing complaint. A 2.4T model most people cannot run, minus the 27B most people could, is a release shaped like a flex rather than a distribution. SourcesAC
Cerebras fell 16% on a number nobody was watching Tuesday
The after-hours drop we carried yesterday settled at 15.7%, to $221.01, and the driver is now clear: an adjusted loss of $2.98 a share against the $0.18 loss expected, an order of magnitude past consensus, on a quarter whose revenue beat and whose full-year guidance was raised. Bloomberg adds the structural read: the hardware segment declined, in what it calls a sign of lumpy demand, while cloud and services grew 287%. Investors are pricing the wafer-scale business on the shape of its transition to services, and a record top line did not change their minds. SourcesBB
LTX-2.5 generates ten seconds of video, with audio, in under seven
Lightricks' LTX released LTX-2.5 open weights on Monday: a 22-billion-parameter asymmetric dual-stream diffusion transformer that generates video and audio jointly, handles multi-shot sequences with consistent characters, lighting and voices across cuts in a single pass, and outputs up to 4K HDR. The number that matters is : a 10-second clip in 6.8 seconds on Nvidia GB200-class hardware, which crosses the line where video generation becomes interactive rather than batch. Commercial use is free below $10 million in annual revenue, ComfyUI workflows shipped day one, and the repo passed 57,000 downloads in its first two days. It is the strongest Western answer yet to Alibaba's Wan line in open video. SourcesAB
MiniMax's geo-fence is not holding
The license on MiniMax-H3 excludes the US, EU, UK and South Korea, and the distribution numbers say the exclusion is words on a page: the official repository is at 1.61 million downloads, the ComfyUI community repackage at 10.4 million, and accelerator forks are proliferating, with one Turbo variant alone past 91,000. A geo-fenced license on an open-weight release restrains exactly the users who read licenses. The other Chinese labs will be measuring what the restriction costs MiniMax in adoption before they copy it. SourcesA
A 150-million-parameter model is the week's most-read paper
BDH-CQ, from the team at Pathway behind the Baby Dragon Hatchling architecture, tops Hugging Face's trending papers at 564 upvotes. The claim: a 150-million- parameter recurrent model that updates an internal memory as inputs arrive and iterates in latent space without verbalizing intermediate steps sets a new cost-accuracy frontier on ARC-AGI-1, the abstraction still find hard. If it replicates, it is evidence that reasoning does not have to be rented by the token from a trillion-parameter model, which is why it is being read. The results are the authors' own and nobody has independently reproduced them yet. SourcesA
Applied Materials reports tonight, the last big capex read before Nvidia
After today's close: Applied Materials, consensus at roughly $3.36 to 3.39 of adjusted earnings per share, up about 35% year over year, on revenue near $9 billion. The items that matter for the AI tape: wafer-fab-equipment demand tied to AI, advanced-packaging and HBM-related tool orders, China exposure, and the October-quarter guide. The stock rose into the print. It is the final heavyweight read on semiconductor capital spending before Nvidia's August 26th report closes the month. SourcesBB
Editorial
Twenty prompts
The number to keep from this week is twenty. That is how many prompts A Security's researchers needed to turn three Zoom bugs into a working zero-click exploit chain, using publicly available models, in under a day. Their own description of what that work used to cost: nation-state infrastructure, elite teams, months.
Set that against the Taiwan disclosure. Twelve attack waves, eight sub-agents at a time, 85 government accounts, and the operators got past the models' guardrails by telling them the intrusion was an authorized penetration test. One sentence of pretext. The most expensive component in the safety stack, refusal training, was talked around with a cover story a phone scammer would recognize.
The load-bearing assumption in enterprise security has always been that exploitation is expensive. Skilled attackers are scarce, so defenders triage: patch what a smart adversary would hit first, accept risk everywhere else. Both stories attack the assumption itself. If a working exploit chain costs twenty prompts, exploitation stops being a scarce skill and becomes a compute budget, and triage by attacker skill stops working. Every unpatched CVE within reach of an agent is now on the schedule.
There is a defense-side version of this argument, and OpenAI has been making it: GPT-5.6-Cyber found a real V8 bug and Google fixed it. The models do cut both ways. But the two sides move at different speeds. An attacker needs one chain to work once; a defender needs the whole installed base patched, and patching runs on release cycles. The Zoom timeline makes the point. The fix shipped June 22nd, and the vulnerable clients will take months to drain out of the world.
What defenders still own is the evidence. An agent intrusion leaves traces no human crew leaves: prompts, tool calls, sub-agent trees, the 160MB archive Dream is working from. SAFE, the incident-reporting standard open for comment since Black Hat, is the first serious attempt to make preserving that material standard practice, and it is voluntary. Voluntary reporting works in aviation because the reporter is also the one in the crash. In security the incentives run the other way, toward silence.
Two observations would settle this within two quarters. Watch the median time from CVE publication to public proof-of-concept: if exploitation is now priced in compute, that number compresses, measurably. And watch whether any organization outside a security lab publishes a full agent trace from a real incident under SAFE. If the first happens and the second does not, offense will have industrialized while forensics stayed artisanal, and that is the bad branch. If exploit timelines hold steady into next year, the twenty-prompt demo was a stunt on a soft target, and this column overweighted it. I would take that outcome. I do not expect it.
Vera Lindqvist
Prediction Watch
Where today's news leaves our open calls. Each one links to the full prediction, its reasoning and the exact test that settles it.
More likely now: Anthropic IPOs before OpenAI (Prediction 2026-08-06-F1). We said Anthropic completes a US listing before OpenAI does. Six of its investors now tell the FT they expect an October debut at $2 trillion or more, while reporting this summer has OpenAI weighing 2027. Expectations are not a , so the call firms rather than settles. Settles August 6th 2027.
Supporting evidence: OpenAI does not IPO in September (Prediction F2). We said OpenAI would not complete a September listing. There is still no public prospectus on , and the reporting around Anthropic's October plans describes OpenAI looking at 2027. Settles September 30th 2026.
Supporting evidence: Nvidia beats the $91 billion consensus on August 26th (Prediction F3). We called for Nvidia's quarter to top the Street. Foxconn just reported AI servers above half its revenue with Vera Rubin racks entering mass-production preparation this quarter, and Supermicro guided its next fiscal year to $65 to 72 billion against a $52.5 billion consensus. The assemblers are seeing the demand Nvidia bills for. Settles August 26th 2026.
No change: stays supply-constrained through 2026 (Prediction F5). We said high-bandwidth memory stays the bottleneck all year. Today cut both ways: trade press has Samsung's HBM4 yield near 80%, which adds supply, while Foxconn's CEO named CoWoS packaging rather than memory as the 2027 ceiling. Neither loosens
- Settles December 31st 2026.
Supporting evidence: nobody solves (Prediction T4). We said no architectural fix reaches broad adoption by next August. The Taiwan intrusion got capable agents to run a four-day campaign by telling them it was an authorized penetration test. Instruction-level manipulation is still the way in, now at state-target scale. Settles August 6th 2027.
New call: DeepSeek holds the August 16th price increase (Prediction 2026-08-13-B1). DeepSeek has published API increases of up to 1,100%, effective Saturday. We put it at 0.70 that the new schedule is still in force in mid-November. The lab that set the industry's price floor raising prices while demand overflows its capacity is the cleanest test yet of whether open-weight competition disciplines API pricing; T3 runs the same experiment on Anthropic from the premium end. Settles November 14th 2026.
China and open weights: a stealth flagship. DeepSeek shipped V4-Pro-0813 with MIT weights a day after Alibaba's Qwen3.8-Max, the heaviest 48 hours for Chinese open weights this year. The promised Qwen3.8-27B has not appeared, Z.ai's rumored GLM-5.5 remains unreleased with GLM-5.2 still listed as current, and the largest open release on record stands at 2.4 trillion , below the 2.8 trillion Prediction T5 needs.
What did not happen. Nothing settled today. GLM-5.5 did not ship. No enforcement action has named a company since the GPAI obligations went live August 2nd. Speaker Johnson has scheduled no hearings. California's suspense-file results land this afternoon.
Sources
- DeepSeek-V4-Pro-0813 — Hugging Face A
- China's DeepSeek upgrades V4 Pro, targets Claude Fable — Decrypt B
- DeepSeek V4-Pro-0813 — Simon Willison C
- DeepSeek API pricing — DeepSeek docs A
- DeepSeek API price hike: V4-Flash compute costs — Android Headlines C
- Near-autonomous AI agents attack Taiwan's nuclear safety agency — The Register B
- China-linked AI agent cyberattack on Taiwan — CNN B
- Anthropic IPO: investors expect $2 trillion October listing — Fortune B
- Anthropic could seek $2 trillion valuation in record IPO — PYMNTS C
- xAI release notes — x.ai A
- SpaceXAI debuts Grok 4.6 — VentureBeat B
- SpaceXAI releases Grok 4.6 — MarkTechPost C
- CoreWeave Q2 earnings, AI demand — CNBC B
- Stock market today, August 12th — The Motley Fool B
- Strong AI cloud demand sends Nebius surging 34% — TradingKey C
- Foxconn second quarter 2026 results — Foxconn A
- Taiwan's Foxconn reports 35% rise in Q2 profit — Business Standard B
- Foxconn server rack operating profit and revenue — DigiTimes B
- Supermicro's stock soars on crushed earnings — SiliconANGLE B
- Super Micro Q4 2026 earnings — 24/7 Wall St C
- One billion monthly users for the Gemini app — Google A
- Google's Gemini app surges to one billion users — TechCrunch B
- Lovable confirms new $13.3B valuation, raises another $400M — TechCrunch B
- AI coding startup Lovable raises $400 million at $13.3 billion valuation — Bloomberg B
- Zoomsday — A Security A
- Zoomsday vulnerability let anyone in a Zoom meeting take over anybody else — Tom's Hardware B
- AI supply chain breach: 2,500 companies, 434,000 CI/CD pipelines — CloudSEK A
- Supply chain attack exposes 2,500 companies — CX Today C
- KOSPI surges for fourth straight day on chip rally — Korea JoongAng Daily B
- Samsung's HBM4 yield reportedly hits 80% — TrendForce B
- Inside the Google executive moves that led to its big AI reshuffle — Reuters via Yahoo B
- IBM, Together AI ink $240 million deal — BNN Bloomberg B
- IBM and Together AI's $240M Nvidia inference cluster — TNW B
- Bank of America to deploy $250 billion — Yahoo Finance B
- Bank of America's $250B infrastructure pledge bets on power — TechTimes C
- Qwen/Qwen3.8-2.4T-A95B — Hugging Face A
- AI News: Qwen 3.8 Max 2.4T and 27B — Latent Space C
- Cerebras (CBRS) Q2 earnings report 2026 — CNBC B
- Cerebras hardware business declines in sign of lumpy demand — Bloomberg B
- Lightricks — Hugging Face A
- LTX-2.5 generates a 10-second AI video in 6.8 seconds — VentureBeat B
- LTX-2.5 open weights release — ComfyUI Wiki C
- MiniMax-H3 — Hugging Face A
- BDH-CQ: In-Context Learning with Recurrent Latent Reasoning — arXiv A
- Open Secure AI Alliance contributions — Nvidia A
- Open-source security groups draft AI agent incident reporting — Axios B
- SAFE guidelines for sharing AI incident data — SecurityWeek C
- Applied Materials stock enjoying AI optimism before earnings — Schaeffer's B
- Applied Materials Q3 earnings preview — TradingKey C
Catch-up: a black-box standard for agent incidents has been open for comment since Black Hat
Catch-up, August 4th, missed by prior issues: the Open Secure AI Alliance, with the Linux Foundation, published , the Shared AI Findings Exchange, as a request for comments at Black Hat: a voluntary standard for reporting AI-agent security incidents, drafted by Cisco, CrowdStrike, Hugging Face, Nvidia and Red Hat. Members would preserve prompts, agent traces, tool calls, identities and credentials from incidents including near-misses, notify directly affected organizations immediately and exposed customers within 72 hours. The model is explicitly aviation accident investigation. This week's Taiwan disclosure is the argument for it: Dream's 160MB evidence archive is the only reason anyone knows what those agents did, and nothing currently obliges anyone to keep such an archive. SourcesAB