An open benchmark of 49 language models on grid logic puzzles shows a steep capability gradient: 85% on 5x5, 20% on 15x15, and most failing a 20x20 hard set.
A single developer published an open test of 49 language models on the same grid logic puzzles. On the Nonobench standard set, solve rates fell from 85% on 5x5 grids to 46% on 10x10 and 20% on 15x15 at each model's best reasoning effort.
Ten 20x20 puzzles with single solutions, five of which line logic alone cannot crack: GPT-6 Astra cleared the standard set, and Claude Opus 5.5 solved 8 of 10. Eleven of fifteen models in that band solved none.
One attempt per puzzle, with 95% confidence intervals reported, and a protocol switch in hard mode. Hard mode shifts from a single 400-character grid string to an array of 20 row strings, because models lose track when forced to compress the whole grid into one string. The code is MIT-licensed on GitHub, routed through OpenRouter and pinned to each lab's endpoint, and the puzzle set is the existing Moyà-Alcover nonograms dataset (CC BY 4.0). Built by Maurice Kleine.
It is a narrow task, and the result does not generalize to "AI can't reason." It does show, on a task a human can finish with pen and paper, exactly where current tools hold and where they break, and the gradient is public, reproducible, and open to anyone who wants to run the next test.