{"id": "2026-08-06-T1", "made": "2026-08-06", "category": "tech", "title": "OpenAI's Astra maths results survive independent scrutiny", "claim": "OpenAI's Astra mathematics results survive independent scrutiny without a material retraction or correction of any of the ten headline results.", "why": "The Lean 4 certificates make outright falsity unlikely. A machine-checked proof is hard to argue with. The live risk here is provenance, not correctness: how the results were produced, not whether they hold. That is what T2 tests.", "confidence": 0.8, "resolves": "2026-12-31", "plain": "Correct if none of the ten headline results is retracted, and no credible published critique shows one is false. Wrong if any is retracted or disproved.", "criterion": "RESOLVES CORRECT if no headline result is retracted, and no credible published critique (arXiv, journal, or a Fields-level mathematician's public statement) establishes that a headline result is false. RESOLVES WRONG if any of the ten is retracted or shown false.", "status": "open", "source": "06-editorial-tech.md", "sources": [{"tier": "A", "title": "Ten advances in mathematics and theoretical computer science — OpenAI", "url": "https://openai.com/index/ten-advances-in-mathematics/"}, {"tier": "A", "title": "An OpenAI model has disproved a central conjecture in discrete geometry — OpenAI", "url": "https://openai.com/index/model-disproves-discrete-geometry-conjecture/"}, {"tier": "B", "title": "An AI math breakthrough sparks calls for new guardrails — Science News", "url": "https://www.sciencenews.org/article/ai-guardrails-erdos-math-problem"}]}
{"id": "2026-08-06-T2", "made": "2026-08-06", "category": "tech", "title": "OpenAI never publishes the raw Astra reasoning traces", "claim": "OpenAI does not release the raw, unedited model reasoning traces behind the Astra results.", "why": "The gap between 'results verified' and 'process verified' is the honest criticism nobody is making. Checking a proof tells you it holds; it tells you nothing about how it was found, or what else was tried.", "confidence": 0.75, "resolves": "2027-02-28", "plain": "Wrong if OpenAI publishes unedited traces for even one headline result. Edited traces, summaries and 'representative excerpts' do not count.", "criterion": "RESOLVES WRONG if OpenAI publishes unedited traces for at least one headline result. RESOLVES CORRECT otherwise. Edited traces, summaries, or 'representative excerpts' do not count.", "status": "open", "source": "06-editorial-tech.md", "sources": [{"tier": "A", "title": "Ten advances in mathematics and theoretical computer science — OpenAI", "url": "https://openai.com/index/ten-advances-in-mathematics/"}, {"tier": "C", "title": "Ten advances in mathematics — Simon Willison", "url": "https://simonwillison.net/2026/Aug/1/ten-advances-in-mathematics/"}]}
{"id": "2026-08-06-T3", "made": "2026-08-06", "category": "tech", "title": "Anthropic holds the Sonnet 5 price increase", "claim": "Anthropic does not reverse or re-discount the Claude Sonnet 5 price increase within 90 days of it taking effect on September 1st 2026.", "why": "The cleanest natural experiment available on frontier pricing power. If a lab can raise the price of its workhorse model and hold it, the race-to-zero story is wrong, and every model of this industry's margins needs redoing.", "confidence": 0.7, "resolves": "2026-11-30", "plain": "Wrong if list pricing drops below $3/$15 per million tokens, or a broad discount restoring pre-September economics appears, before November 30th. Targeted enterprise discounts do not count.", "criterion": "RESOLVES WRONG if list pricing for Sonnet 5 drops below $3/$15 per million tokens, or a broad promotional discount restoring pre-September economics is announced, before 30 Nov 2026. Targeted enterprise discounts do not count.", "status": "open", "source": "06-editorial-tech.md", "sources": [{"tier": "A", "title": "Introducing Claude Sonnet 5 — Anthropic", "url": "https://www.anthropic.com/news/claude-sonnet-5"}]}
{"id": "2026-08-06-T4", "made": "2026-08-06", "category": "tech", "title": "Nobody solves prompt injection", "claim": "No robust general architectural defense against prompt injection achieves broad industry adoption.", "why": "High confidence, because this is an architectural problem rather than a backlog item. Instructions and data share one channel. Until that changes, every defense is a filter. And a filter between a capable search process and its goal becomes part of the search space.", "confidence": 0.85, "resolves": "2027-08-06", "plain": "Wrong if a defense becomes the default at three of the five big labs and independent red-teaming shows better than 95% mitigation on a recognized benchmark. Filters and classifiers alone do not count.", "criterion": "RESOLVES WRONG if a defense is adopted as default by at least three of {OpenAI, Anthropic, Google, Microsoft, Meta} AND independent red-teaming reports >95% mitigation on a recognized benchmark. Filtering/classifier layers alone do not count.", "status": "open", "source": "06-editorial-tech.md", "sources": [{"tier": "B", "title": "Prompt injection still drives most agentic AI security failures in production — Help Net Security", "url": "https://www.helpnetsecurity.com/2026/06/11/owasp-prompt-injection-ai-security-failures/"}, {"tier": "B", "title": "When prompts become shells: RCE vulnerabilities in AI agent frameworks — Microsoft Security", "url": "https://www.microsoft.com/en-us/security/blog/2026/05/07/prompts-become-shells-rce-vulnerabilities-ai-agent-frameworks/"}, {"tier": "B", "title": "OWASP GenAI Exploit Round-up Report Q1 2026", "url": "https://genai.owasp.org/2026/04/14/owasp-genai-exploit-round-up-report-q1-2026/"}]}
{"id": "2026-08-06-T5", "made": "2026-08-06", "category": "tech", "title": "A Chinese lab open-weights a model at 2.8T parameters or larger", "claim": "A Chinese lab publishes open weights for a model at or above Kimi K3's 2.8T parameter scale.", "why": "Cadence-based rather than insight-based. The trend line is steep, consistent, and nobody involved has an incentive to stop. The way this fails is a policy shock, not a technical one.", "confidence": 0.8, "resolves": "2027-02-28", "plain": "Correct if any China-based lab publishes downloadable weights for a model of 2.8T total parameters or more. An announcement without downloadable weights does not count.", "criterion": "RESOLVES CORRECT if any China-based lab publishes downloadable weights for a model with >=2.8T total parameters. Announcement without downloadable weights does not count.", "status": "open", "source": "06-editorial-tech.md", "sources": [{"tier": "C", "title": "China's Open-Weight Takeover — Chris Zeoli, Data Gravity", "url": "https://www.datagravity.dev/p/chinas-open-weight-takeover"}, {"tier": "C", "title": "China's open-weight AI race in 2026: who leads and why", "url": "https://robofutur.com/en/articles/chinese-open-ai-race-2026/"}, {"tier": "C", "title": "Best Chinese AI Model 2026: DeepSeek vs Qwen vs Kimi vs GLM vs MiniMax", "url": "https://checkaimodels.com/en/articles/china-ai-models-landscape-2026/"}]}
{"id": "2026-08-06-B1", "made": "2026-08-06", "category": "business", "title": "Two more lab spin-outs of the Discovery Loop shape", "claim": "At least two additional senior-researcher spin-outs of the Discovery Loop shape occur from major AI labs.", "why": "Scale is rentable now. When the compute is a purchase order rather than a moat, incumbency stops holding people, and the reason to stay becomes equity rather than capability.", "confidence": 0.75, "resolves": "2027-08-06", "plain": "Correct if at least two new ventures are founded by researchers leaving Google/DeepMind, OpenAI, Anthropic, Meta or Microsoft, each with a founder of VP or Distinguished Scientist stature, each raising $100M or more. Acquihires and internal reorgs do not count.", "criterion": "RESOLVES CORRECT if >=2 new ventures are founded by researchers departing {Google/DeepMind, OpenAI, Anthropic, Meta, Microsoft}, each with >=1 founder of VP/Distinguished-Scientist level or equivalent public stature, each raising >=$100M. Acquihires and internal reorgs do not count.", "status": "open", "source": "07-editorial-business.md", "sources": [{"tier": "B", "title": "Jeff Dean and other top AI researchers are leaving Google to launch their own startup — TechCrunch", "url": "https://techcrunch.com/2026/08/05/jeff-dean-and-other-top-ai-researchers-are-leaving-google-to-launch-their-own-startup/"}, {"tier": "B", "title": "The startup idea that convinced a UW computer science legend to leave Google — GeekWire", "url": "https://www.geekwire.com/2026/the-startup-idea-that-convinced-a-uw-computer-science-legend-to-leave-google-after-27-years/"}, {"tier": "C", "title": "Top AI researchers leave Google DeepMind for OpenAI, Anthropic — Crypto Briefing", "url": "https://cryptobriefing.com/top-ai-researchers-leave-google-deepmind-for-openai-anthropic-amid-competition/"}]}
{"id": "2026-08-06-B2", "made": "2026-08-06", "category": "business", "title": "Palantir hits its raised FY2026 guidance", "claim": "Palantir meets or exceeds its raised FY2026 revenue guidance of $8.15B.", "why": "They raised guidance by about $1B mid-year, which implies visibility most software companies never have. The risk is timing rather than demand: government contracts slipping a quarter would do it.", "confidence": 0.8, "resolves": "2027-03-31", "plain": "Correct if reported FY2026 revenue comes in at $8.15B or above.", "criterion": "RESOLVES CORRECT if reported FY2026 revenue >= $8.150B in the Q4/FY2026 report. Direct check against a single reported number.", "status": "open", "source": "07-editorial-business.md", "sources": [{"tier": "B", "title": "Palantir crushes earnings as U.S. AI demand sends revenue soaring 93% — Fortune", "url": "https://fortune.com/2026/08/03/palantir-earnings-guidance-beat-revenue-profit-ai-demand/"}]}
{"id": "2026-08-06-B3", "made": "2026-08-06", "category": "business", "title": "A major breach names prompt injection as a root cause", "claim": "A major publicly-disclosed enterprise security incident occurs with prompt injection or agent compromise as a named root cause.", "why": "340% year-on-year growth in attacks, no architectural fix, and agents shipping into production anyway. This is arithmetic rather than prophecy. The uncertainty is disclosure, not occurrence. It will happen; the question is whether anyone says so out loud.", "confidence": 0.75, "resolves": "2027-08-06", "plain": "Correct if a company above $1B revenue, or a government agency, publicly names prompt injection or agent compromise as a root cause of a breach, and at least one credible outlet covers it.", "criterion": "RESOLVES CORRECT if a company with >$1B revenue, or a government agency, publicly discloses a breach naming prompt injection or autonomous-agent compromise as a root cause, covered by at least one tier-B outlet.", "status": "open", "source": "07-editorial-business.md", "sources": [{"tier": "B", "title": "Prompt injection still drives most agentic AI security failures in production — Help Net Security", "url": "https://www.helpnetsecurity.com/2026/06/11/owasp-prompt-injection-ai-security-failures/"}, {"tier": "B", "title": "OWASP GenAI Exploit Round-up Report Q1 2026", "url": "https://genai.owasp.org/2026/04/14/owasp-genai-exploit-round-up-report-q1-2026/"}, {"tier": "C", "title": "awesome-ai-agent-attacks — timeline of real AI agent security incidents (GitHub)", "url": "https://github.com/webpro255/awesome-ai-agent-attacks"}]}
{"id": "2026-08-06-B4", "made": "2026-08-06", "category": "business", "title": "A Fortune 500 publicly pulls back from an agent program", "claim": "At least one Fortune 500 company publicly announces scaling back or cancelling a major agentic AI program.", "why": "Gartner forecasts over 40% cancellation. That number is unfalsifiable as stated, so this is the same claim rebuilt into a form that can actually be checked.", "confidence": 0.7, "resolves": "2027-08-06", "plain": "Correct if a Fortune 500 company says on an earnings call, in a filing, or in an official statement that it is cancelling or materially cutting an agentic deployment. Anonymous survey data does not count: it has to be attributable.", "criterion": "RESOLVES CORRECT if a Fortune 500 company states on an earnings call, in a filing, or via official statement that it is cancelling or materially reducing an agentic AI deployment. Anonymous survey data does not count; attribution is required.", "status": "open", "source": "07-editorial-business.md", "sources": [{"tier": "C", "title": "AI Agent Adoption 2026: 120+ Enterprise Data Points — Digital Applied", "url": "https://www.digitalapplied.com/blog/ai-agent-adoption-2026-enterprise-data-points"}, {"tier": "C", "title": "Agentic AI Statistics 2026: Adoption, ROI, and Market Size — Unico Connect", "url": "https://unicoconnect.com/blogs/agentic-ai-statistics-2026"}]}
{"id": "2026-08-06-B5", "made": "2026-08-06", "category": "business", "title": "Entry-level developer employment does not recover", "claim": "US entry-level software developer employment (ages 22-25) does not recover to its 2024 level.", "why": "My least confident call, and logged deliberately at 0.60. The interest-rate and over-hiring confound is genuinely hard to separate from an AI effect. At this confidence a miss costs me little and a hit proves little, which is the honest position, not a hedge.", "confidence": 0.6, "resolves": "2027-12-31", "plain": "Wrong if a recognized dataset shows employment for developers aged 22–25 back at 2024 levels. Correct otherwise.", "criterion": "RESOLVES WRONG if a recognized dataset (BLS, ADP, or an equivalent widely-cited series) shows employment for developers aged 22-25 returning to 2024 levels. RESOLVES CORRECT otherwise.", "status": "open", "source": "07-editorial-business.md", "sources": [{"tier": "B", "title": "AI job cuts are rising, but experts say layoffs are only part of the story — CBS News", "url": "https://www.cbsnews.com/news/ai-layoffs-hiring-entry-level-workers/"}, {"tier": "C", "title": "AI Job Displacement Statistics 2026 — AIExposure", "url": "https://www.aiexposure.org/ai-job-displacement-statistics"}]}
{"id": "2026-08-06-F1", "made": "2026-08-06", "category": "finance", "title": "Anthropic IPOs before OpenAI", "claim": "Anthropic completes its IPO before OpenAI completes one.", "why": "Secondary-market bid/ask imbalance plus better revenue quality: Anthropic's revenue is more enterprise and more contracted, which is materially easier to underwrite.", "confidence": 0.75, "resolves": "2027-08-06", "plain": "Correct if Anthropic shares begin public trading before OpenAI's do. Void if neither lists by the resolution date.", "criterion": "RESOLVES CORRECT if Anthropic shares begin public trading before OpenAI's do. RESOLVES WRONG if OpenAI lists first. VOID if neither lists by the resolution date.", "status": "open", "source": "08-editorial-finance.md", "sources": [{"tier": "B", "title": "Anthropic Files Confidential S-1: Joins $3 Trillion AI IPO Race — Yahoo Finance", "url": "https://finance.yahoo.com/markets/stocks/articles/anthropic-files-confidential-1-joins-161008569.html"}, {"tier": "B", "title": "Anthropic Files For IPO, Looking to Beat OpenAI to the Punch — Futurum", "url": "https://futurumgroup.com/insights/anthropic-files-for-ipo-looking-to-beat-openai-to-the-punch/"}]}
{"id": "2026-08-06-F2", "made": "2026-08-06", "category": "finance", "title": "OpenAI does not IPO in September", "claim": "OpenAI does not complete an IPO during September 2026.", "why": "Reported internal lean toward 2027, and $600M of unsold secondary inventory says the private market has not cleared at the current mark. You do not go public into that.", "confidence": 0.8, "resolves": "2026-09-30", "plain": "Wrong if OpenAI shares begin trading on or before September 30th 2026.", "criterion": "RESOLVES WRONG if OpenAI shares begin public trading on or before 30 Sept 2026. RESOLVES CORRECT otherwise.", "status": "open", "source": "08-editorial-finance.md", "sources": [{"tier": "B", "title": "OpenAI files confidential S-1 with SEC — Fox Business", "url": "https://www.foxbusiness.com/markets/openai-signals-potential-stock-market-debut-while-weighing-private-company-advantages"}, {"tier": "B", "title": "OpenAI Targets An IPO As Soon As September At Up To $850 Billion — TechTimes", "url": "https://www.techtimes.com/articles/317955/20260607/openai-targets-ipo-soon-september-850-billion.htm"}]}
{"id": "2026-08-06-F3", "made": "2026-08-06", "category": "finance", "title": "Nvidia beats the $91B consensus on August 26th", "claim": "Nvidia's Q2 FY27 revenue, reported August 26th 2026, exceeds the ~$91B consensus.", "why": "Nvidia has beaten consistently, and this demand is contracted rather than forecast. Confidence is only moderate because a memory-driven supply miss is a live risk, see F5.", "confidence": 0.75, "resolves": "2026-08-31", "plain": "Correct if reported Q2 FY27 revenue comes in above $91.0B.", "criterion": "RESOLVES CORRECT if reported revenue > $91.0B. Direct check against the earnings release.", "status": "open", "source": "08-editorial-finance.md", "sources": [{"tier": "B", "title": "History Says This Is What Will Happen to Nvidia Stock After Aug. 26 — Motley Fool", "url": "https://www.fool.com/investing/2026/08/06/history-says-this-is-what-will-happen-to-nvidia-st/"}, {"tier": "B", "title": "Nvidia stock closes at record, market cap past $5 trillion — CNBC", "url": "https://www.cnbc.com/2026/04/24/nvidia-stock-closes-at-record-pushing-market-cap-past-5-trillion.html"}, {"tier": "C", "title": "NVIDIA market capitalization — CompaniesMarketCap", "url": "https://companiesmarketcap.com/nvidia/marketcap/"}]}
{"id": "2026-08-06-F4", "made": "2026-08-06", "category": "finance", "title": "Credit cracks before equity does", "claim": "Credit spreads on AI-adjacent issuers widen materially before any major AI equity index suffers a 20% drawdown.", "why": "The core structural call. Credit is the sensor and equity is the symptom: this build-out began funded by cash flows and is now increasingly funded by debt, and debt reprices before sentiment does.", "confidence": 0.7, "resolves": "2027-08-06", "plain": "Correct if AI-adjacent credit spreads widen 150bp or more from their August 2026 level before a major AI equity index falls 20% from its high. Wrong if the equity drawdown comes first. Void if neither happens.", "criterion": "RESOLVES CORRECT if a recognized measure of AI-adjacent/data-center IG or HY spreads widens >=150bp from its Aug 2026 level BEFORE a major AI equity index (e.g. SOX, or a recognized AI index) falls 20% from its high. RESOLVES WRONG if the equity drawdown comes first. VOID if neither occurs.", "status": "open", "source": "08-editorial-finance.md", "sources": [{"tier": "A", "title": "Financing the AI boom: from cash flows to debt — BIS Bulletin 120", "url": "https://www.bis.org/publ/bisbull120.pdf"}, {"tier": "B", "title": "Bond Investors Push Back As AI Debt Heads Toward $570 Billion — Forbes", "url": "https://www.forbes.com/sites/robertszczerba/2026/07/17/bond-investors-push-back-as-ai-debt-heads-toward-570-billion/"}, {"tier": "B", "title": "The $3 Trillion AI Data Center Build-Out Becomes All-Consuming For Debt Markets", "url": "https://energynow.com/2026/02/the-3-trillion-ai-data-center-build-out-becomes-all-consuming-for-debt-markets/"}]}
{"id": "2026-08-06-F5", "made": "2026-08-06", "category": "finance", "title": "HBM stays supply-constrained through 2026", "claim": "HBM memory remains supply-constrained through the end of 2026.", "why": "Underpins both the memory thesis and the memory-bound scaling argument. Capacity takes years to build and the orders are already placed, so the supply side cannot surprise inside the window.", "confidence": 0.75, "resolves": "2026-12-31", "plain": "Wrong if two or more of Samsung, SK Hynix and Micron state on an earnings call that HBM supply has moved into balance or oversupply.", "criterion": "RESOLVES WRONG if two or more of {Samsung, SK Hynix, Micron} publicly state on an earnings call that HBM supply has moved into balance or oversupply. RESOLVES CORRECT otherwise.", "status": "open", "source": "08-editorial-finance.md"}
{"id": "2026-08-06-F6", "made": "2026-08-06", "category": "finance", "title": "OpenAI signs another $10B+ non-Azure compute deal", "claim": "OpenAI announces at least one additional non-Azure cloud or compute commitment of $10B or more.", "why": "Tests the Microsoft-concentration thesis from OpenAI's side. Diversifying compute is the tell that the relationship is a supplier arrangement rather than a partnership.", "confidence": 0.65, "resolves": "2027-08-06", "plain": "Correct if OpenAI publicly announces a new compute or cloud agreement of $10B or more with a provider other than Microsoft Azure. Expansions of the existing Amazon deal count only if separately announced at that size.", "criterion": "RESOLVES CORRECT if OpenAI publicly announces a new compute/cloud agreement >= $10B with a provider other than Microsoft Azure. Expansions of the existing Amazon deal count only if separately announced and >= $10B.", "status": "open", "source": "08-editorial-finance.md", "sources": [{"tier": "B", "title": "AI Circular Deals: How Microsoft, OpenAI and Nvidia Keep Paying Each Other — Bloomberg", "url": "https://www.bloomberg.com/graphics/2026-ai-circular-deals/"}, {"tier": "B", "title": "Microsoft demand backlog doubles to $625 billion — Fortune", "url": "https://fortune.com/2026/01/28/microsoft-stock-drops-azure-growth-slows-capex-spending-q2/"}]}
{"id": "2026-08-11-F1", "made": "2026-08-11", "category": "finance", "title": "Nvidia's compute-financing MOUs produce $100B of closed vehicles in a year", "claim": "By August 11th 2027, definitive financing vehicles, funds or facilities established under the Nvidia partnerships with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR total at least $100B in publicly announced committed capital.", "why": "The $500B figure is a target attached to MOUs, and MOU-to-close attrition in infrastructure finance is real. But six institutions with $4T of AUM do not lend their names to a press release they intend to strand, and the demand side (labs converting capex to leases) is visibly forming the same week. One fifth of the headline inside a year is the honest over-under between marketing and a market.", "confidence": 0.65, "resolves": "2027-08-11", "plain": "Wrong if a year from the announcement the six platforms have publicly closed less than $100B in definitive, committed financing vehicles. Restated MOUs, targets and ambitions do not count; announced closings and committed funds do.", "criterion": "RESOLVES WRONG if by 2027-08-11 public announcements from Nvidia or the six named partners of definitive (non-MOU) vehicles, funds or credit facilities under these compute-financing platforms sum to less than $100B in committed capital, per company press releases or major financial press. RESOLVES CORRECT otherwise.", "status": "open", "source": "reports/daily/2026-08-11.md"}
