Four quantum error correction code families on a common benchmarking pipeline: three on IBM gate based hardware and a fourth digitized GKP branch (a continuous variable code) on a continuous variable simulator.
On a 56-qubit IBM gate-model chip, a new preprint asked a quantum error-correction decoder to find a single planted error across 4096 circuit runs. The decoder was supposed to light up only where the error was. Instead, it lit up across most of the lattice, and almost none of those activations pointed back to the injected target. The authors call the result plumbing, not a threshold: an audit of the syndrome-to-decoder interface on real hardware, not a step toward a useful machine.
The paper benchmarks four code families on the same kind of hardware pipeline. Three of them ran on IBM gate-model circuits: a 5-data-qubit repetition code, a distance-5 rotated surface-code Z-check extraction layer, and the Z-check half of the Steane code standing in for a compact CSS-LDPC benchmark (a class of quantum error-correcting codes built from classical low-density parity-check designs). The fourth branch is a digitized-GKP sampling pipeline, where 'GKP' is shorthand for the Gottesman-Kitaev-Preskill code, a continuous-variable encoding that approximates ideal logical qubits using grid states. It used finite-squeezed Gaussian continuous-variable readout and injected q-shifts, binned into the same Z-check interface, and was run on a PennyLane Gaussian simulator rather than a chip. Each branch was driven by 4096 shots, with clean and injected streams, using LiDMaS+ request construction. The plotted baseline is minimum-weight perfect matching (MWPM); union-find and hard-decision belief-propagation / min-sum are replayed for interface validation rather than to declare a winner.
On the repetition and CSS-LDPC hardware branches, the dominant expected syndrome and correction survived every injected target. On the 56-qubit routed surface circuit, the picture flipped: exact localization of the planted error collapsed to between 0.003 and 0.108, while broader 'target-containing' localization held at 0.279 to 0.642. About half of the decoders' activations had nothing to do with the planted error, a gap the paper attributes to hardware-induced syndrome activation drowning the signal rather than to any flaw in the decoder. The digitized-GKP branch sat in between, with exact q-shift localization at 0.350 to 0.495 and target-containing localization at 0.417 to 0.608, a tighter range but one that lives entirely inside a PennyLane Gaussian simulator.
Correction-volume panels in the preprint track the mean minimum-weight correction weight per decoded stream, giving the auditable interface a concrete number to replay. The authors position the work as supporting a reproducible, comparable syndrome-to-decoder pipeline that future cross-code experiments can lean on, paired with a same-day companion paper, arXiv:2607.19446, which extends the comparison across quantum software stacks. Both are preprints and neither makes a fault-tolerant claim. The honest watch item is whether the syndrome-noise gap the IBM surface run surfaces can be narrowed enough that a single threshold number across code families stops being premature.