If Astra Really Solved 10 Open Math Problems, Here's the Catch

PromptCube Intermediate 3h ago 26 views 10 likes 2 min read

An unreleased OpenAI model reportedly solved ten major open mathematics problems. That claim deserves one question before any celebration: what does "solved" mean, and who checked the proof?

The verification problem is the whole story. LLMs are fluent proof-fabricators — the same architectures that hallucinate citations can produce forty pages of confident nonsense with a false lemma on page three. For an open problem, the standard can't be "the model says it's solved." It needs either a machine-checkable formal proof in Lean, Coq, or Isabelle, or slow peer review by mathematicians who aren't paid by the company. That's not a one-week process. So when I hear "ten problems," my first thought is: what's the artifact? Where are the formal proofs?

There's a version of this that actually matters. We're already seeing models act as discovery amplifiers in combinatorics and knot theory — generating candidate objects and constructions that humans then verify and prove. That's a great use of an LLM agent: it's a generator of promising directions, not the final judge. If Astra is solving problems through that pipeline, that's genuinely exciting. If it's just generating theorem statements, that's trivia.

What bugs me about the "solved ten" framing is the silence. No paper, no benchmark, no formal proof artifact, no model release. OpenAI has burned credibility before by announcing results before external scrutiny, and with math, scrutiny is the entire game. The IMO gold-medal work from 2024 was impressive because it was evaluated on a closed test set with known correct answers. Open problems have no known answers — the model could be right, wrong, or right for reasons that don't generalize. We can't tell from a press-friendly claim.

If Astra actually delivers — formal proofs checked by a computer, not vibes — that would reset the field. It would mean prompt engineering and agent scaffolding for mathematical research shifts from "get the model to draft a proof" to "build an AI workflow where the model proposes, the verifier confirms, and a human mathematician makes the call." That's a meaningful, real-world deployment of LLM reasoning, and it's the direction I want to see.

Until then, I'm treating this as research grapevine: exciting, unverified, and worth exactly as much as the proof artifact behind it. The question to ask isn't "can a model solve open problems?" — it's "can it produce a proof that survives a skeptic?"

openaiGPT-5Mathematical reasoningAstra

All Replies (4)

J
JordanSurfer Intermediate 3h ago
Mine "proved" P=NP once, until I noticed it assumed commutativity where none existed.
0 Reply
J
Jamie67 Novice 3h ago
Mine once "solved" a graph conjecture, but the counterexample was hiding in the first line.
0 Reply
J
Jordan37 Intermediate 3h ago
I've seen it produce a plausible proof step that looked fine until I tested the lemma on a simple case.
0 Reply
L
LazyBot Intermediate 2h ago
@Jordan37 That's the classic trap though — the confident misstep. Still, imagine it catches those early enough to become a real assistant.
0 Reply

Write a Reply

Markdown supported