IBM, Q CTRL (a quantum control software vendor), and QESEM (a quantum error mitigation tool from startup Qedma) ran the same five workloads on one 156 qubit IBM Heron r3 quantum processor.
On a single 156-qubit IBM Heron r3 processor nicknamed "Pittsburgh," three competing error-management stacks ran the same five workloads back-to-back for the first time. The new cross-stack benchmark, posted to arXiv in August, does not crown a winner. It names a two-axis choice that practitioners using near-term quantum machines will have to make: which errors hurt most, and which fixes cost how much in billable processor time.
The three stacks are IBM's own Qiskit Runtime (raw execution plus its built-in measurement twirling and TREX+twirling mitigation), Q-CTRL Performance Management (an external error-suppression suite), and QESEM, Qedma's error-mitigation engine. All three ran on the same physical hardware: one IBM Heron r3 chip, in one vendor software snapshot, on the same day.
The workloads split into two camps. The Sampler side tested four standard circuits used to stress-test near-term machines: Bernstein-Vazirani (a hidden-bit search), quantum phase estimation, GHZ-state preparation (a maximally entangled multi-qubit state used to verify that qubits still "talk" to each other), and randomized mirror circuits, run at up to 100 measured qubits. The Estimator side tested chain-averaged magnetization and correlation observables (the average spin and pairwise spin-spin coupling) of an eight-layer transverse-field Ising circuit (a canonical many-body physics model used to stress-test quantum simulators) at 25, 50, and 75 qubits, scored against an exact matrix-product-state reference (a classical-computation method that gives the exact answer for one-dimensional quantum systems within a controlled approximation).
Across six Ising observable and system-size cases, aggregate mean absolute error came in at 0.0883 for IBM raw execution, 0.0807 for IBM TREX+twirling, 0.0285 for Q-CTRL, and 0.0188 for QESEM. Read a different way: relative to raw execution, Q-CTRL and QESEM reduced aggregate error by 3.10x and 4.70x, respectively.
QESEM, the most accurate of the three, used between 7.5x and 11.1x the reported QPU time of Q-CTRL. QPU time here means billable processor runtime on the shared quantum machine. Q-CTRL held its own on the structured Sampler workloads, posting the best results on three of them, while keeping reported QPU times in the same order of magnitude as the IBM configurations.
The authors frame full quantum error correction as still too costly for routine use, and position suppression and mitigation as the practical near-term lever. They are not claiming any tool here will scale to a fault-tolerant machine. The result is a snapshot on one chip, in one vendor software version, on one day.
The choice is straightforward. If the team cares more about the size of the wrong answer than the size of the bill, QESEM's 4.7x error reduction is the lever. If it cares more about the size of the bill, Q-CTRL's 3.1x reduction at near-baseline QPU time is the lever. The paper's IBM baselines matter too: raw execution is the floor, and TREX+twirling is a free, vendor-native option that already trims roughly 9% of aggregate error on this benchmark.
Generalizing the result to other qubit counts, topologies, or hardware families (IonQ, Quantinuum, Rigetti) is not supported: the benchmark ran on one IBM Heron r3 instance. The paper is a preprint, not peer-reviewed, and result attribution to the stack versus device calibration drift on a single hardware instance should be hedged. Replication on a second Heron r3 instance, or on a different chip generation, would tighten the picture.
Three commercial vendors allowed an apples-to-apples test on identical hardware, identical workloads, and a disclosed methodology. The next preprint that repeats the run on a different Heron r3 instance, or on the next chip generation, will say whether the 3.1x to 4.7x accuracy gap is a property of the software stacks or of one machine on one day.