FIELD NOTES
The mathematicians are having a week.
OpenAI claims to have solved a version of the Navier–Stokes problem, one of seven famous, highly difficult mathematical challenges defined by the Millennium Prize. I remember them walking past the Millennium Prize challenges, hung up in the main halls of MIT. They almost served like an everyday reminder of the pursuit of greatness.
To see problems that once felt like monuments to human ambition potentially fall so quickly… and amid questions about how the work was done… is an odd thing to sit with.
In response, influential mathematician Terence Tao joined 24 fellow medallists in warning that the race to solve problems is losing sight of what mathematics is for. Getting an answer and developing understanding are different achievements.
I’m drawing parallels with Google DeepMind’s Atlas project, which this week published predictions for nine billion possible DNA changes. We can now generate an extraordinary number of hypotheses about biology. Drug discovery candidates seem imminent. Scale can get us surprisingly far.
We may not understand how it actually works.
Does that matter?
The AI-pilled will argue that humans are already lapsing far behind AI, and our human lack of true understanding is what will become. Others argue that these solves are superficial and doesn’t truly advance the field.
That gap is also becoming harder to ignore inside the labs. Anthropic researcher Jacob Coxon resigned this week, warning that frontier lab competition was taking precedence over safety. The labs “are racing straight to self-improving superintelligence and gambling with our lives”. A few days later, Anthropic and Open AI and DeepMind leaders subsequently called to slow down capability development of AI.
Does that matter? Or are we already too late?
I’m not convinced we can put technology progress back on hold. I’m also not sure I’d want to, if it means delaying something that could save lives.
But the possibility of curing a disease can’t be the answer to every concern about what else these systems might do.
Maybe we’re too late to stop the race. I hope we’re not too late to change how it’s being run.
-Tara
THE DOWNLOAD
OpenAI Claims a Navier–Stokes Solution. Mathematicians Question the Race to Solve Problems.
OpenAI released an argument and computer-checkable formalization claiming finite-time blowup for a forced version of the Navier–Stokes equations, which describe fluid motion.
The company says the effort involved roughly 10,000 concurrent agents, 88 hours of research and another 17 hours of formalization in Lean, a language that allows a computer to check the logical steps of a proof. The claim concerns specific forced formulations of the problem; its scope and acceptance still require mathematical scrutiny. (OpenAI)
On Friday, Terence Tao, one of the most influential mathematicians working today, joined 24 other Fields medallists in publishing a declaration about AI’s growing role in mathematics. They acknowledge that models can solve major problems. Their objection is that treating those problems as benchmarks can separate the answer from the understanding mathematics is meant to develop.
A result normally leads to talks, simplifications, new methods and connections to earlier work. The declaration argues that rushed announcements can short-circuit that process and raise attribution questions. It is not a technical refutation of OpenAI’s proof. (Terence Tao)
Why it matters: Last week’s Fermat formalization showed machines taking on more of the work of checking mathematics. This week raises the next question: once machines can produce results and formalize them, who makes those results understandable and useful? A proof can be correct without yet giving other researchers a method they can build on.
Meta Launches Muse, Joining the Race for a Personal AI Agent
Meta launched Muse, a personal AI agent designed to take on tasks, remember context and work toward longer-term goals. People can message it through a dedicated app or WhatsApp, give it access to services and let it continue working in the background. (Meta)
The ambition is an assistant that knows enough about your life to do more than respond to individual requests. Planning a trip might involve coordinating dates, comparing options and following up as circumstances change. The useful part is continuity: you should not have to explain everything again each time.
Instinct, a buzzy personal-agent startup that raised $250 million at a $2.5 billion valuation in weeks while still in private beta, is pursuing a similar idea. Its assistant connects to applications and devices, with a stated focus on understanding personal context and following through on everyday tasks. (Instinct)
Early experience shows both the appeal and the difficulty. An Atlantic reporter describes successful purchasing and coordination tasks, alongside reports of unwanted reservations and excessive booking requests. Meta has built a separate permission system, Sentinel, to control which actions Muse can take or send back for approval. These safeguards remain to be proven. (The Atlantic / Meta’s technical account)
Why it matters: Personal agents become useful through accumulated context and repeated delegation.The product challenge is building an assistant people trust with progressively more of their lives.
DeepMind Maps the Predicted Effects of Nine Billion DNA Changes
Google DeepMind released AlphaGenome Atlas, a resource containing predictions for approximately nine billion possible single-letter changes in human DNA.
A DNA change can affect more than the protein a gene produces. It can also alter when a gene is active or how its instructions are processed. The Atlas gives researchers a way to look up predicted effects across a large set of possible changes, helping them decide which variants to investigate experimentally. (DeepMind)
The new development is the resource itself: a broad, precomputed map that laboratories can consult rather than generating every prediction separately.
There is an important limit. These are nine billion predictions, not nine billion experiments. Checks of selected variants do not validate the entire map, and single-letter changes do not cover every kind of genetic variation.
Why it matters: Models are starting to become reference infrastructure for science. A shared prediction map could save researchers time and help prioritize scarce experimental resources. But it will also influence which questions get investigated. Knowing where the predictions fail becomes part of using the atlas well.
Robots Are Getting Better at Learning From Us
Showing a robot how to do something is harder than it sounds. A human demonstration has to become instructions that make sense for a different body, with different joints, sensors and ways of handling objects.
Several papers this week tackle parts of that translation.
SEED-UMI uses a shared exoskeleton interface to bring human demonstrations closer to the movements a robot can execute. The authors report 70% success across five tasks and threefold faster data collection. The aim is to preserve more of the useful information in a demonstration as it passes from person to machine. (Paper)
Show-Harness, from Singapore’s NUS ShowLab, approaches the problem through instructions. It gives language models a constrained set of actions that a control system translates into robot movements. Its experiments cover ten pick-and-place task combinations on Franka and AgileX hardware. The model chooses an action without having to generate every motor command itself. (Paper)
A third study shows how much a robot can learn through immediate feedback. It demonstrates dexterous pen writing by rapidly estimating how movements affect the pen, reporting roughly 18 seconds of initialization on a laptop CPU and 0.6 mm mean in-plane writing error. (Paper)
Why it matters: Better demonstrations and interfaces can reduce how much a robot must discover from scratch. Part of the work is making human intent easier for the machine to use; another part is letting it correct its movements as it goes. Both can make useful behavior less dependent on enormous training runs.
Anthropic Says Moonshot Served Claude Responses to Kimi Users
Anthropic alleges that Moonshot, the company behind Kimi, forwarded some customer requests to Claude and displayed the responses as its own.
The report describes almost 300,000 customer requests over one ten-day period, mostly routed to Opus through a network of 5,380 accounts. Anthropic separately alleges that Moonshot collected exchanges for training. (Anthropic)
That is more specific than the usual allegation of training on a competitor’s outputs. It concerns which provider actually answered a customer’s request.
It also falls short of establishing that “Kimi K3 is Claude.” Routing requests, learning from another model’s outputs and having identical underlying models are different claims. The report does not establish that every Kimi request was forwarded.
These remain allegations from a competitor. Bloomberg reports Anthropic’s account rather than independently verifying the traffic attribution. (Bloomberg)
ALSO:
Only 5% of Surveyed Chip Designers Reported First-Silicon Success. Siemens EDA and Wilson Research Group report a decline from 14.4% in 2024 among IC/ASIC respondents. This measures reported success on the first silicon iteration, not manufacturing yield. It is a useful counterpoint to faster design cycles: getting a design ready for fabrication and getting production-ready silicon remain different milestones. (Siemens)
Google Signs a 22-Year Nuclear Power Agreement in Finland. The Fortum contract begins in 2028 and reaches 50% of Loviisa’s capacity for 2030–2049. Separately, the US Department of Energy closed financing of up to $1.9 billion for the 615 MW Duane Arnold restart. Both support existing nuclear assets; neither should be counted as newly operating generation. (Fortum / DOE)
Sparse Models Can Save Compute and Still Need More Distinct Data. Mixture-of-experts models activate only part of their parameters for each token. A new study finds that repeatedly recycling training examples can erode their advantage over dense models. Regularization helps in the experiments, so this identifies a tradeoff rather than a universal limit: cheaper computation does not automatically make repeated data equally useful. (Paper)




