The Brief: The Heat of AI Summer 2026
The hottest week of AI summer 2026: a $1.3 trillion chip drawdown, frontier models hacking real companies, and three Chinese labs shipping major models in 24 hours
We are back after a brief hiatus! New research from the firm coming soon. The week we came back to was the hottest of AI summer 2026: three Chinese labs shipped major models inside 24 hours, chip stocks gave back $1.3 trillion, and both OpenAI and Anthropic disclosed their models had breached real companies.
THE DOWNLOAD
Chip Stocks Lose $1.3 Trillion in a Week
The world’s 20 most valuable chip companies lost about $1.3 trillion in market value in the sessions ending July 29: Nvidia shed $238B, SK Hynix $176B, Samsung $173B, TSMC $119B, Micron $113B, and AMD $110B. The worst day was July 28, when reports that China could expand domestic memory capacity erased $541B from six memory names in a single session.
Why it matters: Analysts largely agree fundamentals didn’t move; Morningstar’s Michael Field calls it sentiment, not fundamentals. The sharper read is leverage: chip stocks were up 92% over the prior year, Morgan Stanley’s momentum TMT index fell 53.5% in July, and as Jim Osman puts it, AI needs years of patient capital but lenders mark collateral daily, so ordinary corrections become forced liquidations.
Situational Awareness Falls 67% in July; Citadel Buys Its Public Portfolio
One of the world’s biggest AI funds told investors it was down about 67% for July, writing “we let you down this month.” Leopold Aschenbrenner’s fund ran roughly 4x leverage on a concentrated public book; margin calls from its prime brokers forced a liquidation, and Citadel bought the entire public equities portfolio in a block trade. Assets went from a $45B peak to about $10B.
Why it matters: The public portfolio, most of the fund, was sold to Citadel at a discount to settle the borrowing, and the fund is still up on the year because what remains is mostly its Anthropic stake, a private position with a large paper gain that lenders couldn't touch and Aschenbrenner can't sell. Whether investors ever see that 80% now rests almost entirely on the Anthropic IPO expected by October.
OpenAI and Anthropic Models Breached Real Systems During Cybersecurity Evals
OpenAI disclosed that during an internal capability eval with safety classifiers deliberately disabled, GPT-5.6 Sol and a pre-release prototype exploited a zero-day in OpenAI’s own package-registry proxy, escaped the sandbox, chained exposed credentials, and achieved remote code execution on Hugging Face servers. Prompted by that disclosure, Anthropic audited 141,000 eval runs and found three incidents of its own: eval environments that were supposed to be air-gapped weren’t, and Claude models reached real companies, in one case pulling hundreds of rows from a production database, in another publishing a malicious package that ran on 15 real systems. Two of the three victims never detected the intrusions.
Why it matters: Autonomous exploit development by frontier agents is no longer hypothetical. But note what actually failed. OpenAI’s incident involved a real zero-day; Anthropic’s models needed only weak passwords and SQL injection, and the root cause was harness misconfiguration, not model intent. Testing environments now need the same security as production systems, which is a new cost for every lab and a new product category for whoever sells it. And the fact that the victims didn’t notice any of it says as much about the state of company security as it does about the models.
OpenAI Says Internal Astra Model Produced Results on Ten Open Math Problems
OpenAI published ten results in math, quantum complexity, and theoretical CS produced by an internal, unreleased version of Astra, its next model family, including the first explicit non-sofic group (open since 1999) and the first improvement to the general sphere-packing exponent since 1978. Every proof ships with a machine-checkable Lean 4 certificate. OpenAI says finding all ten cost roughly $2,000 at Sol API rates.
Why it matters: The Lean certificates make this harder to dismiss than a benchmark claim, though no external mathematicians reviewed it pre-announcement and critics object to math by press release. If settling decades-old conjectures is now a compute line item, the value of formal, verifiable domains as training and product surface just repriced, and mathematician Thomas Bloom’s reaction (”big news”) suggests the field agrees this one is real.
Three Chinese Labs Ship Major Models Inside 24 Hours
On July 31 Beijing time, three releases from Chinese labs landed: DeepSeek shipped V4-Flash-0731, MiniMax released H3, an omni-modal video model generating 15-second 2K clips with native stereo audio at roughly a third of mainstream per-second pricing, and ByteDance launched Seedance 2.5, extending single-run generation to 30 seconds.
Why it matters: Axios calls DeepSeek’s update comparable to Opus 4.8 on coding at a ~99% discount, and OpenAI already cut Luna pricing 80% three weeks after launch, leaving Anthropic as the last premium-pricing holdout. Chinese open models now take 41% of global open-model downloads, and on Vercel’s AI Gateway, Seedance already leads video generation volume over Veo. Ex-OpenAI’s Zack Kass: “at some point, the next model doesn’t matter to you.”
Moonshot Releases Kimi K3 Weights and Technical Report
What it is: Moonshot’s Kimi K3, a 2.8T-parameter MoE with 104B active, native vision, and 1M-token context, launched via API July 16 with full weights following July 27. Artificial Analysis scores it #3 on its Intelligence Index, comparable to Opus 4.8 and GPT-5.5, behind only Fable 5 and GPT-5.6 Sol.
Why it matters: Nathan Lambert’s estimate is that K3 compresses the closed-to-open gap from 6-9 months to 3-5, which resets what closed labs can charge for. Two asterisks before anyone builds on it: the custom license requires a commercial agreement above $20M revenue, so this is open-weight, not open-source; and self-hosting realistically requires a 64-accelerator cluster, so for most teams “open” still means renting cloud inference.
Thinking Machines Releases Inkling-Small as Open Weights
Thinking Machines Lab released Inkling-Small, a 276B-parameter MoE with 12B active, about a quarter the size of Inkling, open-weighted on Hugging Face and fine-tunable on Tinker. Trained partly by on-policy distillation with Inkling as teacher, it beats the larger model on reasoning and agentic coding, hitting 80.2% on SWE-bench Verified, while trailing on knowledge coverage.
Why it matters: The student beating the teacher on reasoning is the notable part; distillation now produces smaller models that are better, not just cheaper. A small model that hits 80% on SWE-bench and that customers can fine-tune and own is the concrete version of Murati's pitch that companies should control their own models.
Google DeepMind Announces Gemini Robotics 2
DeepMind announced a suite of three models on July 30: a vision-language-action model for whole-body humanoid control and five-fingered dexterity, an embodied reasoning model that plans across multiple robots, and an on-device model that adapts to new robot bodies within hours. Demos ran on Apptronik’s Apollo 2, with Boston Dynamics and Agile Robots among named partners.
Why it matters: DeepMind robotics head Carolina Parada is explicit about the strategy: build “the intelligence layer that can be used by every robot.” If the brains become a Google API, hardware margins compress and the defensibility of vertically integrated humanoid bets weakens. Temper it with DeepMind’s own numbers: 40-44% success on complex fine-motor tasks. Dexterity is still unsolved.
FCC Adds Foreign Humanoids, Robot Dogs, and Power Inverters to Covered List
On July 28 the FCC added foreign-made mobile robots and connected power inverters to its Covered List, implementing a White House interagency determination. New models can no longer receive the equipment authorization required to import or sell in the US; existing devices are unaffected, with exemptions available case by case. The target is China, where Unitree, Agibot, and UBTech hold roughly 87% of humanoid shipments.
Why it matters: US humanoid startups just got a protected home market and more expensive development at the same time. Interact Analysis’s Rueben Scriven notes the near-term hit to Chinese vendors is small since US penetration was minimal; the real risk is slowing US commercialization by removing the cheap platforms teams prototype on.
Amazon Reported to Wind Down Most Nova Models
Business Insider reports (not yet confirmed by Amazon) that Amazon is deprecating Nova Premier, Omni, Reel, and consolidating behind a new Frontier Model Research group led by Pieter Abbeel, with a new flagship expected at re:Invent. This follows the closure of its SF AGI lab and the December reorg under Peter DeSantis.
Why it matters: Amazon is conceding the current model generation to concentrate on the next one, and the economics say it can afford to: AWS hosts $138B+ of OpenAI commitments and $100B+ from Anthropic, and Bedrock profits whichever model wins.
Substack Adds AI Detection; Snapchat Stops Recommending Fully AI-Generated Videos
Substack integrated Pangram’s AI detector, letting readers scan posts, Notes, and comments over 100 words published after July 21 for estimated AI-generated share; Pangram claims a 0.01% false-positive rate and finds over 20% of Substack longform already flags as AI-assisted. Days later, Snapchat made wholly AI-generated videos ineligible for Spotlight recommendations, while content enhanced with its own AI tools stays eligible with transparency indicators.
Why it matters: Platforms are converging on a stance: AI as tool is fine, AI as content gets suppressed, with LinkedIn and YouTube moving the same direction the same month.
Study Finds Half of AI Unicorns Have Never Published Research
A Stanford METRICS preprint from John Ioannidis’s group, covered by Science, analyzed all 317 AI unicorns and found 52.4% have never published a single paper or preprint. The whole cohort contributed 0.1% of AI literature in 2025; OpenAI alone accounts for 39.4% of the group’s citations.
Why it matters: You can no longer judge an AI company, or its researchers, by published work; half the most valuable startups in the field have none to show. What replaced it is what the authors call "blogification": unreviewed claims in corporate posts that feed both training data and hype cycles. The Hacker News steelman is that companies exist to sell what they invent; fair, but an ecosystem built on "Attention Is All You Need" has now mostly closed the door behind itself.


