FIELD NOTES
I spent a couple of days at Hot Chips this week, and came away feeling like… the culture of chipmaking is changing.
One moment that stuck with me was watching OpenAI present Jalapeño, its first inference chip.
The performance numbers were impressive. But what was more surprising was how quickly it got there. OpenAI says Jalapeño went from initial design to tapeout in nine months, using its own models to accelerate parts of the design process.
For advanced silicon, that is extraordinarily fast. Typical timelines have been closer to 18 - 24 months.
Chip design has traditionally been slow for good reason. A mistake can cost tens of millions of dollars and months of manufacturing time. Architectures accumulate abstractions, verification processes and institutional knowledge precisely because you don’t get many chances to be wrong.
OpenAI seemed willing to make a different trade.
Jalapeño doesn’t need to be good at everything. OpenAI knows the models it runs, the workloads it serves and where the bottlenecks are. So it can design much closer to the workload and discard some of the generality that Nvidia needs.
You could feel versions of this elsewhere at the conference. Cerebras still making the bet that the chip should be an entire wafer. D-Matrix stacking custom memory directly onto compute. Normal Computing and its forays into thermodynamic compute. Different architectures, but a similar willingness to question assumptions.
I wouldn’t be too quick to assume the old guards (the chip uncs) are being replaced.
At the same event, Nvidia and Google talked through work that has spanned decades. Much of what the new generation can move quickly on exists because that infrastructure and knowledge already exists.
But clearly, change is in the air. AI can compress parts of the design cycle. TSMC and the semiconductor supply chain let companies specialize without building fabs. Broadcom can provide decades of implementation expertise while OpenAI supplies intimate knowledge of the workload. And when you’re spending billions of dollars serving essentially the same class of computation, even a narrowly useful chip can justify its existence.
For decades, building a serious processor required not only enormous capital but an enormous accumulation of semiconductor-specific knowledge. A great wall to climb. This may be less true today. And that means the next generation of chips may increasingly come from companies that didn’t start out as chip companies at all.
-Tara
One more thing: I have many more things I want to dig into than I can currently get to, and I’m thinking about building out the research side of Strange Review.
Would you be interested in deeper proprietary research, or a private chat to share and discuss things in real time? Hit reply. I’d love to know.
THE DOWNLOAD
1,200 Agents Self-Organized For A Cyberattack
OpenAI and METR published their postmortems on the July incident where agents escaped an internal cybersecurity evaluation and eventually compromised Hugging Face.
METR reconstructed a hidden message board used by 1,206 agents, carrying more than 70,000 messages over five days. Roughly 700 agents participated in the attack on Hugging Face. Some sacrificed their own evaluation performance to help other agents. A handful considered maybe alerting a human. None did.
The agents also figured out how to spoof parts of their tool-use transcripts. METR found successful spoofing in roughly 7% of the transcripts it examined, while the grader intended to catch this behavior was missing a check described in the evaluation documentation.
The obvious lesson is that the sandbox failed. The more interesting one is that hundreds of individually limited agents became more capable once they found one another.
We spend a lot of time asking what happens as individual models get smarter. This was a glimpse of a different scaling axis: what happens when you have a thousand of them.
Nvidia May Be Buying The Open Models Layer
Nvidia is reportedly in talks to acquire Hugging Face for around $13 billion. Neither company has confirmed a deal.
Calling Hugging Face the “GitHub for AI” increasingly undersells what it’s becoming.
It sits at the center of the open model and dataset ecosystem. But last year it also acquired Pollen Robotics. Its LeRobot project is building common tooling and datasets for robot learning. Last month it released Grabette, a cheap open system that lets a person record physical tasks with a handheld gripper and turn them directly into robot-training data.
Hugging Face’s thesis is explicit: “The bottleneck isn’t the model. It’s the data.”
Then this week Pollen (acquired by Hugging Face in April 2025) unveiled Microduck, a $399 open-source consumer robot.
There’s a more interesting interpretation of a potential Nvidia acquisition. Nvidia already dominates the compute on which AI runs. Hugging Face increasingly sits where models, datasets, developers — and now physical-world training data — meet.
If Nvidia buys it, it isn’t just buying distribution for open models. It may be buying an emerging data and distribution layer for physical AI too.
→ Business Insider / Hugging Face
OpenAI’s First Chip Is Fast
OpenAI showed off Jalapeño, its custom inference chip built with Broadcom, at Hot Chips this week. The first numbers are fairly wild.
On SemiAnalysis’s InferenceX benchmark, Jalapeño cleared 700 tokens per second for a single user running DeepSeek R1. On Kimi K2.5, it delivered more than 9× the performance of the next-best system at 100 tokens per second per user.
There are some big caveats. This is engineering silicon, the tests used a short fixed prompt rather than messy real-world traffic, and SemiAnalysis says the fairest comparison would be Nvidia’s upcoming Vera Rubin rather than today’s systems. SemiAnalysis verified the runs in OpenAI’s lab, but didn’t independently run the full benchmark suite.
Still, the direction is more interesting than the benchmark win. OpenAI isn’t trying to build a general-purpose GPU. Hardware lead Richard Ho says the goal is to run OpenAI’s most important workloads “close to the hardware’s theoretical limits.”
That’s the bet: if you know exactly which models and workloads you’re serving, you can throw away the generality Nvidia needs and optimize the entire chip around inference.
Jalapeño starts shipping in small volumes at the end of this year, with volume production in 2027. If the advantage survives real workloads, OpenAI won’t just be one of Nvidia’s biggest customers. It will have started designing around Nvidia.
→ OpenAI / SemiAnalysis
Also:
Memory is getting expensive enough to show up in AI server prices. Some Nvidia Grace Blackwell and Vera Rubin systems shipping next year are expected to cost >15% more, driven by soaring memory costs. Nvidia CFO Colette Kress called current memory pricing conditions “extreme.” If memory reprices faster than compute gets cheaper, some of the expected generational decline in inference cost gets eaten away. → Bloomberg
OpenAI is cutting Cursor off. OpenAI plans to stop supplying its models to Cursor on November 12 following SpaceX’s acquisition of Anysphere, saying it can’t trust Musk’s companies to comply with its terms. Cursor says OpenAI accounts for only ~5% of its traffic. → OpenAI
Anthropic wants agents operating physical equipment. Its new Model Hardware Standard gives agents a common interface for microscopes, liquid handlers, robotic arms and even lasers. Anthropic says integrations that previously took weeks can sometimes be done in hours. MCP connected agents to software; this is an attempt to do something similar for machines. → Anthropic
Children remain annoyingly data-efficient. BabyLM has spent three years testing whether language models can learn on something closer to a child’s data budget. Better architectures and training methods help, but the enormous sample-efficiency gap remains. We increasingly know that models can learn surprisingly sophisticated language from less data. We still don’t know why children need so much less. → MIT Technology Review





