OpenAI proved an 80 year old Erdős conjecture; the same month, researchers showed these systems can collapse on simple puzzles, and the gap between them is the story.
In May 2026, OpenAI proved a famous 1946 conjecture of the Hungarian mathematician Paul Erdős in a single shot. The model was a "general-purpose reasoning model," a class of AI systems trained to show their intermediate steps before answering, called large reasoning models or LRMs. The result was research-grade math: a strict lower bound that beat the long-standing "square grid" construction and pushed the unit distance problem past the threshold Erdős had guessed. A month later, the number theorist Will Sawin sharpened that bound on his own, posting an explicit exponent of n^1.014 and cleanly disproving a second Erdős conjecture in the process.
The same month, an Apple Machine Learning Research team published a paper called "The Illusion of Thinking." Their finding, summarized in a Quanta Magazine synthesis of the active debate: chain-of-thought reasoning, the stream of synthetic text a model emits before its final answer, is not what its name suggests. Under surprisingly simple conditions, LRMs show "complete accuracy collapse." A separate study from the Santa Fe Institute, also cited in the Quanta analysis, showed the same class of model crushing carefully designed visual-reasoning benchmarks using what the authors call "surface-level shortcuts," pattern matches that look like reasoning and produce correct answers on the test set, but do not transfer.
Both findings can be true at once. That is the part worth sitting with.
The OpenAI result is the harder of the two to dismiss. Erdős's unit distance problem asks how many pairs of points, at unit distance apart, can fit on an n-by-n grid without sharing a column or row. The best known upper bound is roughly n^(4/3), set by Spencer, Szemerédi, and Trotter. The OpenAI model produced a lower bound whose exponent strictly exceeds 1, and the company released a rewritten chain-of-thought PDF documenting the reasoning trace, including discussion of a sharper n^(1+c/log log n) scale. Sawin's preprint translates the result into a concrete, citable exponent (epsilon greater than 0.014), improving the OpenAI bound by a factor of roughly 10^36, according to one third-party analysis.
The math community's reaction has been measured. The result was filed in a corporate blog post, not in a peer-reviewed journal, and Sawin's note is also an arXiv preprint, not yet refereed. Number theorists are reading the chain of thought, not the press release.
The critique is harder to dismiss for a different reason. The Apple team did not run LRMs on hard math problems; they ran them on puzzles the models should have aced. The collapse happens on tasks that look elementary, where the chain of thought is short enough to inspect. The Santa Fe result is more troubling still: the models win the benchmark. The shortcut is invisible from the scoreboard. If the same dynamic operates on harder problems, a correct answer is no evidence of correct reasoning. It is just a correct answer.
This is the question the AI labs are now living inside. When a model emits a long, plausible chain of thought and lands on a research-grade result, who inside the lab, or outside it, can tell whether the chain was load-bearing or decorative? The chain of thought is the only window the user has. It is also generated, not recorded.
Gary Marcus and Ernest Davis, the AI critics quoted in the Quanta synthesis, put the version of the worry that is hardest to wave off: a gold medal at the International Mathematical Olympiad is so hard that successful mathematicians may "highlight [it] on their CVs all their lives." LRMs have won that medal. The same LRMs fail simple puzzles and pattern-match their way through analogy tests. Both are real. The question is what the medal certifies.
Terence Tao has used AI to rediscover or improve solutions to 67 problems across mathematical analysis, combinatorics, geometry, and number theory. The result of Tao's work is hard to read as shortcut behavior, and hard to read as anything but a serious tool. The contradiction does not resolve there either.
A useful frame for the next time a lab announces a reasoning win: ask which definition of "reasoning" the result is being measured against, and which shortcut could in principle explain it. The OpenAI unit distance proof has a citable chain of thought and a sharpened, named exponent. That is a higher bar than a benchmark win. It is also a bar most announcements do not clear.
Sawin's preprint went up on arXiv in May 2026. The OpenAI model that produced the original bound is not publicly described in enough detail to replicate. The next round of math results will be read against that gap.