The Advisory Group on Mathematics and AI wants AI labs to name the model, show the prompts, show the compute, and skip the marketing. OpenAI's GitHub release is the first test case.
OpenAI published 722 mathematics manuscripts on a public GitHub repository, spanning 372 result families and produced by a frontier model whose name has not been disclosed (The Verge). The company says the average result used the compute equivalent of about three hours of ChatGPT Pro thinking.
That volume is the capability data point. The governance data point landed in the same release window. The Advisory Group on Mathematics and Artificial Intelligence, an independent panel of elite mathematicians convened as an external check on AI-generated proofs, published its first recommendations in late September. The recommendations ask labs to release results promptly, route them through established academic channels where possible, and disclose the model name, the prompts used, and the compute spent. The advisory group also implored AI companies to "refrain from treating the release of mathematical results as marketing vehicles to promote their models," a practice it said inflicts "significant harm on the mathematical community" (The Verge).
"Show your work," in this context, is a four-part checklist. Name the model so other labs can replicate the result. Publish the prompts so reviewers can see what the model was actually asked. Publish the compute so the cost of generating a new proof is auditable. Route the result through a journal or established preprint server so it enters the literature on the same terms as a human-authored paper. OpenAI's openai/math preprints repository passes the first test, a public release, and partly passes the second, by including model reasoning summaries and compute estimates alongside each preprint. It does not yet pass the model-name or prompts tests, which is why this release lands in the same week as the advisory group's recommendations rather than after them.
The advisory group's "hundreds of open problems" estimate frames the scale of the batch. OpenAI's own September statement was more conservative: it said its model had "resolved more than 100 long-standing open problems across most areas of mathematics." The two figures are not contradictory. The repository is larger than the September announcement, and AGMAI's count is a separate scope assessment, not a per-problem verification. Some share of the 372 result families is likely to be minor extensions or re-derivations rather than genuinely closed open problems. That share has not been published.
One preprint in the repository is already attracting scrutiny. Titled "Integer multiplication below n log n" and dated September 23, 2026, the paper claims a bound that is, on its face, tighter than the 2019 O(n log n) result from Harvey and van der Hoeven for integer multiplication (preprint). The bound is worth reading carefully. If it holds, it is a real algorithmic step. If it does not, it is precisely the kind of result the new disclosure norms are meant to surface early. The repository does not run a peer review pass before the model publishes, so the first filter is the math community reading the abstract.
The wider pattern is institutional. AGMAI is a working parallel review track, and its first published recommendations are the test case for whether AI-generated math can enter the existing literature on the same terms as a human-authored paper. The next test is the next release from the same repository. If that release names the model, publishes the prompts, and routes through a journal rather than a vendor-controlled repo, the disclosure norms have weight. If it does not, the standards stay voluntary, and the math community will need to build a way to enforce them.