OpenAI’s Astra Just Cracked 10 Long-Standing Math Problems — Here’s Why That’s Both Historic and Incomplete
OpenAI's internal Astra model reportedly produced ten advances on long-standing mathematics problems, verified via the Lean formal proof checker — but critics note an undisclosed failure denominator and the rapid partial replication by a public Claude model raise important questions about the scope of the claim.
Mathematics has long served as AI’s most comfortable proving ground. Benchmarks are clean, answers are verifiable, and Olympiad medals photograph well in press releases. But on August 3, 2026, OpenAI published something that feels qualitatively different: a claim that its internal model, Astra — its next major release — generated ten genuine advances on long-standing open problems across geometry, cryptography, quantum computing, and pure mathematics.
According to The Neuron, which covered the story in its August 3 issue, some of these results settle conjectures that researchers had been wrestling with for decades. Others improve on previously known theoretical limits. The announcement came alongside a 249-page paper, Lean proof files, and model-written reconstructions of the reasoning behind each result.

What Astra Actually Did
Two results stand out as especially significant. First, Astra reportedly made the first improvement since 1978 to a major high-dimensional sphere-packing limit — a problem about how tightly equal objects can fit together in many dimensions. This is not a classroom exercise; sphere-packing bounds have deep implications in coding theory and information transmission. A gap in the frontier since 1978 is a very long time.
Second, Astra constructed a “non-sofic group” — a mathematical object that researchers had long wondered might even exist. In doing so, it disproved Connes’s rigidity conjecture, resolving a question that had sat open in the mathematical community for years.
The methodology deserves as much attention as the results themselves. Humans worked alongside Astra to prepare each proof, but crucially, the model converted every proof into Lean — a formal proof verification language whose automated checker independently validates each logical step. This is not a matter of trusting the model’s output. When Lean says a proof is correct, it means the machine has verified every inference in the chain. That verification layer is what separates this from previous AI math claims that relied on expert human judgment alone.
OpenAI also published the full reasoning walkthroughs, making the thinking process — not just the conclusions — available for scrutiny.
The Cost Dimension
Here is a detail that puts the scale of computation into perspective: OpenAI estimates that successful solution tokens would cost roughly $2,000 (approximately ₹1,70,000 at current rates) at Sol API rates per result. That figure hints at the enormous search process happening beneath the surface. Astra is not solving these problems the way a human mathematician does, building intuition over years of study. It is exploring a vast idea space with what The Neuron describes as — in a deliberate approximation — “gigawatts of compute.”
This reframes what is actually happening. The model’s contribution is not elegant insight in the human sense. It is industrialised search over mathematical possibility space, guided by humans who select verifiable problems, and confirmed by proof software that checks every answer. The loop looks something like this: humans identify problems with checkable answers, Astra searches aggressively, Lean verifies, and the result is a formally confirmed advance.
If that loop can be made reliable and repeatable, it could compress years of mathematical trial and error into days.

The Sceptics Arrive Quickly
No major AI announcement escapes scrutiny for long, and this one attracted some sharp criticism within hours.
Gary Marcus, writing on his Substack, acknowledged Astra’s achievement but pushed back on the implied scope of the claim. His core argument, as reported by The Neuron: checkable mathematics does not prove universal scientific reasoning. Math is uniquely suited to this kind of AI-driven breakthrough because it offers right-or-wrong feedback and an essentially unlimited supply of synthetic practice problems. Cancer research, military strategy, climate modelling, and the vast majority of real-world decisions do not share those properties. Praising Astra for mathematical advances and then extrapolating to “AI can now do science” is, in Marcus’s framing, a significant overreach.
Then came a more concrete challenge. In a follow-up post, Marcus highlighted the work of Levent Alpöge, a mathematician at Anthropic, who reportedly reproduced roughly half of Astra’s results using public Claude Fable — not a special internal model, not a custom setup, not with internet access — just a generic prompt and full model autonomy, within 24 hours of the announcement.
The Neuron notes this comparison needs an apples-to-apples public evaluation to be conclusive. Astra is not publicly available; Claude Fable is. The problems Alpöge attempted, the prompts used, and the exact criteria for “reproduction” all matter enormously. But even as an incomplete data point, it meaningfully weakens the narrative that Astra represents a singular, unreachable threshold. If a publicly available model can reproduce five of the ten results, the story becomes less about one extraordinary model and more about a general capability shift across the frontier.
The Missing Denominator
The most important number OpenAI has not published is arguably the most important number of all: how many problems did Astra attempt that it did not solve?
Noam Brown, one of the researchers involved, publicly acknowledged that OpenAI tried other major problems without success. But the full count of attempts, the nature of those problems, the human guidance provided at each stage, and the total cost of failed runs — none of that is in the public record. Without a denominator, a ten-out-of-unknown hit rate is impossible to interpret statistically.
The credible next test, as The Neuron frames it, is a pre-committed open-problem set where every failure counts alongside every success. This is standard scientific practice: you declare your hypotheses before you run the experiment. A curated retrospective of wins, however impressive each win is individually, tells you less than you need to know.

What This Actually Changes
None of the above caveats should obscure what is genuinely new here. The verification-via-Lean approach is significant precisely because it sidesteps the trust problem that has plagued AI-generated mathematics. You do not have to believe the model; you just have to run the checker. That changes what human-AI collaboration looks like in research contexts.
The pipeline — humans selecting problems, models searching idea space, proof software confirming results — is now a demonstrated workflow, not a theoretical one. Whether Astra is dramatically more capable than its peers at this specific loop, or whether the loop itself is what matters and any sufficiently powerful frontier model can participate, is an open question. The Alpöge replication suggests the latter deserves serious consideration.
For researchers in India and globally working in fields adjacent to formal mathematics — cryptography, theoretical computer science, quantum algorithms — this is worth watching closely. The tools for participation in this loop are becoming more accessible, and the problems that yield to it may multiply faster than expected.
As The Neuron puts it: AI has not “solved” science, but it may have permanently changed how some science gets done. The academic reviewers tasked with verifying 249-page AI-generated proof documents have, perhaps, earned our sympathy.
The breakthrough may exceed one secret model. Frontier AI appears to be entering a reusable research loop — and that changes the calculus for everyone working at the boundary of what is mathematically known.
The question now is not whether AI can produce mathematical advances. That is established. The question is whether the field will hold these claims to the same standards of reproducibility and transparency it demands of any other scientific result. So far, the answer is only partially yes — and the missing denominator is where the real story lives.
