An internal version of Astra, OpenAI's unreleased AI model, produced raw arguments that human authors turned into preprints; Lean, an open source proof checking system, verified each step.
OpenAI published a bundle of work this week that the company is calling "ten advances in mathematics and theoretical computer science." Each result, OpenAI said in a blog post, was proposed by an internal version of its next model family, Astra, and then checked by a Lean proof certificate before publication. The list spans high-dimensional geometry, coding theory, group theory, operator algebras, and arithmetic complexity, all long-standing open problems.
The mechanism is the news here, more than the specific theorems. In each case, Astra produced the raw argument, human researchers prepared the manuscript, and a Lean verifier, software that mechanically checks a proof's logical steps, produced a certificate. The artifact isn't a peer-reviewed paper. It's a preprint with a machine-checked proof and a company as the claimant. "An internal version of Astra, @OpenAI's next major model family, solved 10 major open problems in mathematics, quantum complexity, and theoretical computer science," wrote Brown. "We believe it will be a major step for scientific reasoning."
The validation chain is the part of this story that ages. The work is real: a Lean certificate is a real artifact, and the ten proofs are published on GitHub with a PDF companion and a separate document of the model's reasoning. But a company choosing which ten results to highlight is not the same as a math community accepting them. Several of the results are major claims: a disproof of Connes's rigidity conjecture in operator algebras, a construction of non-sofic groups, a step toward the Cohn–Elkies bound in sphere packing, and an arithmetic-formula lower bound of order n^4/log n for the permanent. None of these have peer-reviewed validation yet in the bundle. The framing of "ten of ten" is OpenAI's selection, not an external audit.
That selection is worth picking apart. One of the ten is a disproof of a long-standing conjecture, not a positive resolution. Several are new bounds, narrowing the gap on a known result rather than closing it. The list is curated. OpenAI's post describes the work as "a step toward" answers in several cases, accurate, but easily flattened in coverage into "AI solved ten open problems."
OpenAI also says the total tokens needed to find the solutions would cost "roughly $2,000 at Sol API rates." Sol is OpenAI's inference API, so the figure is a token bill on the company's own pricing, not a difficulty benchmark. It does suggest the compute footprint is small. It does not say the problems were easy.
The validation pattern isn't brand new. May. The new ten-result post frames that earlier work as already inspiring follow-on developments, which gives the workflow a second data point within the same model effort.
OpenAI is also pairing the release with a distribution channel: a "ChatGPT for Academic Researchers" initiative giving 100,000 scientists and mathematicians free access to its best models. That's a real move, not just a research output. It positions the preprints as a recruiting and adoption funnel, not only a contribution to the literature.
Readers can use three questions to separate a genuine validation chain from a company-curated preprint: who picked the problems, who wrote the manuscript, and who certified the proof? If the answers are "the lab," "a human at the lab," and "a proof checker that the lab also runs," the result is a preprint with strong tooling but a narrow validation chain. The work can still be good. Lean certificates are real artifacts, and the underlying arguments may be sound, but the gate is the lab's, and the next step is whether the math community treats these preprints as routine or exceptional.
The preprints are out. The Lean certificates are in. The math community, on its own clock, will decide whether the ten are breakthroughs or starting points.