10,000 AI Agents Just Cracked a $1 Million Maths Problem — And Sparked a Fight Between OpenAI and Anthropic
OpenAI claims to have solved the Navier-Stokes Millennium Prize Problem by deploying 10,000 AI agents simultaneously, generating 2.7 million messages and 130 billion tokens over 88 hours. The announcement has sparked a fierce controversy with NYU and Anthropic researchers who were already working on related problems.
Mathematics has seven Millennium Prize Problems — puzzles so fiendishly difficult that the Clay Mathematics Institute attached a $1 million (roughly ₹8.5 crore) reward to each one. For decades, the best human minds on the planet have made only fitful progress on them. According to a report in The Neuron’s September 8, 2026 issue, OpenAI has now claimed to solve one of them: the Navier-Stokes existence and smoothness problem. The method it used, however, is as remarkable as the result — and has ignited a genuine controversy across academia and the AI industry.

What Is the Navier-Stokes Problem, and Why Does It Matter?
The Navier-Stokes equations describe how fluids — water, air, blood, the atmosphere — move through space over time. Engineers and physicists use them constantly, from designing aircraft wings to modelling monsoon patterns. The mathematical problem, however, cuts deeper: do these equations always produce smooth, well-behaved solutions, or can they eventually generate a singularity — a point where velocity becomes infinite, essentially causing the mathematical model to “blow up”?
This is not an abstract curiosity. If the equations can blow up, it suggests there are fundamental limits to how well classical fluid dynamics can describe nature. Mathematicians have wrestled with this question for decades without a decisive answer. OpenAI says its system has now found one, and the answer is that yes, a fluid can develop a singularity in finite time.
The 10,000-Agent Strategy: Parallel Thinking at Scale
The most striking aspect of this claimed breakthrough is not which model produced it — it is how the result was reached. According to The Neuron’s reporting, OpenAI did not assign the problem to a single model and wait for inspiration to strike. Instead, it deployed roughly 10,000 AI agents simultaneously, each equipped with tools, code execution capabilities, and crucially, the ability to share useful discoveries with one another.
The scale of the operation is staggering by any measure:
- The multi-agent system generated 2.7 million agent messages over the course of the effort.
- It produced approximately 130 billion output tokens in total.
- The entire effort concluded in roughly 88 hours.
Think of it less like a lone genius having an insight and more like mobilising an enormous, highly coordinated research institute that never sleeps, never loses focus, and can instantly share partial results across thousands of parallel workstreams. The agents collectively converged on a proposed proof showing that a fluid singularity can form in finite time.
Once a candidate proof emerged, a separate, unreleased model — described by OpenAI as significantly more capable than GPT-6 Astra — stepped in to formalise and verify the proof in Lean, a software tool that checks mathematical arguments step by step with machine-level rigour. This verification step matters enormously, because claimed proofs in mathematics are notoriously easy to get subtly wrong.
The Academic Controversy: Did OpenAI Jump the Queue?
Results of this magnitude rarely arrive without friction, and this case is no exception. The Neuron notes that NYU mathematician Tristan Buckmaster and Anthropic researcher Levent Alpöge were already working on closely related fluid-dynamics problems when OpenAI announced its breakthrough.
OpenAI’s account, published on its website, is that the company heard rumours of progress being made in this space and responded by pointing thousands of agents at the remaining Millennium Problems itself. It explicitly denies having accessed the researchers’ unpublished work and maintains that its proof was produced independently.
Buckmaster and Alpöge see it differently. TechCrunch and WIRED have both reported on the allegations in detail, with TechCrunch’s headline describing the situation as OpenAI having “fought dirty” on a career-making problem. Scientific American, meanwhile, has published analysis of what the broader mathematics community makes of the proof itself — a community that will need to scrutinise the result carefully before any prize is formally awarded.

The ethical questions here are genuinely thorny. Academic researchers often share informal progress updates and preprints within their communities. If a well-resourced AI lab can monitor the direction of frontier research and then sprint past it with thousands of compute-heavy agents, what does that mean for the norms of scientific collaboration? These questions do not have easy answers, and they are likely to resurface as AI systems become more capable.
What This Means for How We Should Judge AI Capability
The Navier-Stokes story reframes a debate that has been building quietly in the AI field: what is the right unit of analysis for measuring what an AI system can actually do?
For most of the current AI era, the default benchmark has been the single-turn interaction — what does one model say in response to one prompt? Leaderboards, evaluations, and public perception have all been shaped by this frame. The Navier-Stokes effort suggests that frame is becoming obsolete.
As The Neuron’s analysis puts it: start judging frontier AI by what a system can accomplish with enough agents, tools, time, and compute — not just what one chatbot answers in one conversation. The Navier-Stokes result was not achieved by a smarter model; it was achieved by a smarter deployment strategy combined with an enormous amount of parallelism.
For practitioners and organisations thinking about how to apply AI, this is a meaningful shift in mental model. The question is no longer only “which model is best?” but also “how do you architect a system of agents to tackle a problem that is too large, too long, or too multifaceted for any single model interaction?”
Three Implications Worth Watching
- 1. Compute as a research resource. The Navier-Stokes effort consumed 130 billion output tokens. At current market rates, that represents a significant expenditure that only a handful of organisations globally could absorb. Agentic research at this scale is currently a capability of the very largest labs, not individual researchers or even most enterprises.
- 2. Formal verification becomes essential. The use of Lean to verify the proof is not a footnote — it is arguably the most important part of the story. As AI systems generate increasingly complex outputs, human review alone will not be sufficient to validate them. Automated formal verification is the mechanism by which trust in AI-generated mathematics (and potentially AI-generated code, regulatory filings, and scientific claims) can be established.
- 3. The remaining Millennium Problems are now targets. OpenAI says it turned its agents toward all the remaining Millennium Problems once it heard rumours of progress. Six problems remain unsolved. Whether this specific result holds up under mathematical scrutiny or not, the episode signals that the most capable AI labs are now treating century-old unsolved problems as tractable engineering challenges rather than permanent mysteries.
The Bigger Picture: A Kasparov Moment for Mathematics?
Scientific American’s coverage draws a comparison to Deep Blue defeating Garry Kasparov in chess — a moment that did not end chess but permanently changed how humanity understood the relationship between human and machine cognition. Whether the Navier-Stokes claim survives peer review or not, something similar may be happening in mathematics.

The mathematics community will need months, perhaps longer, to fully assess the proof. Formal verification in Lean reduces but does not eliminate the possibility of error — the formalisation process itself must be checked. And the controversy surrounding how the result was obtained will need to be addressed transparently if the broader scientific community is to trust and build on it.
What is already clear, regardless of how the peer-review process concludes, is that the architecture of AI problem-solving has taken a qualitative step forward. Ten thousand agents, 2.7 million messages, 130 billion tokens, 88 hours — these are the new coordinates of frontier AI capability. The era of the lone chatbot is giving way to something that looks far more like a coordinated, always-on research organisation. India’s technology sector, its mathematics departments, and its AI policy bodies would do well to study this development closely, because the implications — for research, for intellectual property norms, and for the future of scientific discovery — are only beginning to unfold.
**Source:** The Neuron newsletter, issue dated September 8, 2026.
