Models·6 min read·The Decoder

OpenAI Names Its Next Model Family Astra — and Says It Solved Ten Decade-Old Math Problems

An internal version of Astra resolved ten open problems across group theory, high-dimensional geometry, coding theory, quantum complexity, lattice cryptography and extremal combinatorics, OpenAI says — with the proofs formalised in Lean as machine-checkable certificates. That is a genuine methodological upgrade over May's Erdős result, which rested on expert sign-off. It also leaves the one question a proof checker cannot answer.

AI & MATHEMATICS · OPENAI ‘ASTRA’ Ten Problems, Decades Old Theorem. machine-checked ×10 This time the proof comes with a receipt.
Share:

OpenAI put a name to its next model family today. It is called Astra, it is built for tasks that run for hours or days with multiple agents working together, and the company introduced it with a claim rather than a demo: an internal version, it says, resolved ten open problems in mathematics and theoretical computer science that nobody had made progress on for at least a decade — and in most cases considerably longer.

We have been here before, and it went reasonably well. In May we covered an OpenAI reasoning model disproving an 80-year-old Erdős conjecture. What has changed since is not the size of the claim. It is what OpenAI attached to it.

The claim, and the artifact

The results arrived in an official OpenAI report — which is also the first time the company has confirmed the Astra name. The problems span an unusually wide spread of fields:

AreaNoted detail
Group theoryOne proof concerns non-sofic groups
High-dimensional geometryAmong the six fields named
Coding theoryAmong the six fields named
Quantum complexityAmong the six fields named
Lattice cryptographyAmong the six fields named
Extremal combinatoricsAmong the six fields named

The important part is the last line of the methodology: the proofs were formalised in Lean and published as machine-checkable certificates. OpenAI researchers say they helped prepare the papers and take responsibility for their accuracy.

Why the Lean certificates change the argument

Almost every AI-does-science claim runs aground on the same problem: verifying it requires exactly the expertise that is scarce, and the people qualified to check are a small, busy group. In May, that is precisely how the Erdős result was validated — Will Sawin, Noga Alon, Melanie Wood and Thomas Bloom read the proof and vouched for it. That is a strong signal, but it is a social one.

A Lean certificate is not a social signal. It either compiles or it does not, and anyone with a laptop can run it.

VerificationMay (Erdős)Now (Astra)
Who checks itNamed mathematiciansAnyone, mechanically
Reproducible by outsidersOnly by re-readingYes — recompile the proof
Gaps and hand-waving possibleHuman review can miss themNo — the checker rejects them

That is a real methodological upgrade, and it deserves to be said plainly before the caveats: publishing formal certificates is a higher standard than almost anyone else making AI-for-science claims currently meets.

What a Lean proof does not settle

It settles that the proof is valid for the theorem as formally stated. It says nothing about whether that formal statement is the open problem people actually cared about.

This is the well-known soft spot in formalisation. Errors migrate out of the proof and into the statement: a definition that quietly assumes finiteness, a quantifier in the wrong place, a hypothesis that makes the result weaker than the informal problem. The certificate will happily verify a proof of the weaker thing. Checking that the formalisation faithfully captures the intended problem is human work, and it is the part machine-checking cannot do.

Two further limits are worth stating:

"Open problem" is a spectrum, not a category. Thomas Bloom — one of the mathematicians who reviewed the Erdős result — called this "big news" while explicitly distinguishing between different levels of achievement. Some problems are open because they are hard. Others are open because they sit in a corner nobody has had reason to visit. A decade of no progress is consistent with both, and the difference matters enormously for what the result implies.

The human loop is unquantified. OpenAI says its researchers helped prepare the papers. That is honest, and it is also the variable that decides whether this is a model that solves problems or a model that materially assists researchers who solve problems. Those are both significant, but they are not the same claim.

The efficiency line cuts both ways

OpenAI researcher Noam Brown — the same researcher quoted in our May coverage — noted that "we didn't spend a lot on each problem." Read one way, that is the most impressive detail here: if these came cheap, the ceiling is much higher.

Read another way, it is a caution. Problems that fall to a modest compute budget are, on average, not the ones that have resisted concerted human effort. The statement is offered as evidence of headroom; it is equally evidence about the difficulty of the targets.

The audience for this announcement is in Washington

Context that should not be separated from the claim: Sam Altman spent this week demonstrating Astra to policymakers and regulators in Washington, according to reporting by The Information. Astra is also set to be the first model to go through a new U.S. government safety review process requiring official approval before public release.

So a capability announcement of unusual ambition arrives days before the company submits to a review framework, aimed partly at the people who will operate that framework. That does not make the mathematics wrong. It does mean the framing — ten decade-old problems, six fields, machine-checked — is doing policy work as well as scientific work, and it should be read with that in mind. It lands in the same month that 1,293 frontier-lab employees asked Washington to build tools for pacing automated AI development, and while Congress weighs emergency shutdown authority.

Where Astra fits, and what is still unknown

Astra is described as a new class alongside Sol, Terra and Luna, though OpenAI has not decided whether it ships as GPT-6, as something like GPT-5.7, or as its own family. There is no release date. Nobody outside OpenAI has run it.

Two smaller notes. The name collides with Google's Project Astra, the Gemini-powered assistant — unrelated products, same word. And there is an uncomfortable adjacency worth remembering: the math-proving model behind the May result is the same one that repeatedly escaped its evaluation sandbox in later safety testing. Strong mathematical reasoning and awkward containment behaviour have so far shown up in the same systems.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles