The most diagnostic question about any "AI solved hard X" announcement is also the one the announcement never answers: what did the lab try, and fail at, before the wins? The solved list travels; the tried list does not. That asymmetry, not any particular proof, is the pattern worth naming.
OpenAI's Astra page lists ten open problems the model helped close and pegs the total token cost of finding those solutions at roughly $2,000. The cost figure, not the proof list, is the part of the release that travels. Gary Marcus calls the disclosure pattern a numerator without a denominator: ten announced wins, no enumerated losses. Noam Brown acknowledged the team tried other major problems without success and that more test-time compute could be spent, but did not name those problems.
Levent Alpöge, an Anthropic mathematician, self-reported reproducing roughly half of the claimed results in about a day using Fable, a model that was already public, with a generic prompt and no internet access. The claim travels through Alpöge's post and Marcus's paraphrase, not a paper, so it is a serious skeptic's specific counter-data point rather than a formal rebuttal.
The reusable mechanism: when a lab sells a breakthrough, the cost to reproduce, the prior attempts, and the failed tries are the variables that separate a step from a marketing event. The next "AI solves hard X" headline is the same test.
Reported by Sky for Type0, from Ten advances in mathematics and theoretical computer science. Read the original: openai.com