The LuGo algorithm moves classical setup work off a quantum chip and onto the Frontier supercomputer, cutting a benchmark fluid simulation circuit from ~2 million gates to ~91,000.
Oak Ridge National Laboratory researchers have cut the work a quantum chip must perform to model viscous fluid flow by more than 95%, by routing the heavy classical setup to a Frontier supercomputer node and leaving the quantum chip with a 91,000-gate circuit where it once needed nearly 2,000,000.
The work, peer-reviewed in Future Generation Computer Systems, Volume 178, 2026 and honored with a 2026 R&D 100 Award, is named LuGo. The pattern it embodies is hybrid: a supercomputer does the bookkeeping, a quantum chip does only the arithmetic only it can perform.
Each "gate" in that count is a basic operation a quantum chip performs, and running millions of them in sequence is what makes quantum fluid simulation impractical on current hardware. Quantum chips are small and noisy. LuGo does not build a better quantum chip. It shrinks the work the chip has to do by moving the parts a classical computer already handles well onto a single Frontier node (the Department of Energy's flagship exascale machine, which peaks at about 2 exaflops; LuGo uses one node).
The benchmark is a textbook viscous-flow problem: fluid squeezed between two parallel plates, a setup physicists call Hele-Shaw flow. Solving it requires inverting a linear system, the math done by an algorithm named HHL, after its inventors Harrow, Hassidim, and Lloyd. HHL is the part of quantum computing that promises speedups for linear algebra, but it has been stuck for years because the quantum circuit grows huge, fast. LuGo's contribution is a preprocessor that uses classical high-performance computing to prepare the quantum state with far fewer operations, leaving the quantum chip to run only the phase-estimation step at the algorithm's core.
The arXiv preprint walks through the mechanism: most of the gate budget in the original HHL routine comes from classical initialization, not the quantum phase estimation. The preprocessing ported to Frontier collapses that initialization by exploiting structure in the Hele-Shaw problem. The result, in the preprint's headline benchmark, is the roughly 2,000,000-to-91,000 gate reduction.
The team then validated the compressed circuit across five different quantum processors, a step that turns a paper algorithm into cross-vendor plumbing. According to the Oak Ridge Leadership Computing Facility release, tests ran on Quantinuum's H-1 trapped-ion machine, IBM's Marrakesh and Sherbrooke superconducting chips, and IQM's Garnet and Sirius superconducting systems, coordinated through DOE's Quantum Computing User Program and its Quantum User Expansion for Science and Technology (QUEST) effort. The work is also benchmarked on NERSC's Perlmutter system.
The 95% number is a circuit-depth reduction, not a wall-clock speedup. Today's quantum hardware cannot run these circuits at problem sizes that beat classical CFD solvers outright, and the HHL approach to linear systems still has a long-standing readout problem: getting a useful answer out of the quantum state, not putting the right state in, is its own unsolved engineering task. ORNL frames LuGo as a preprocessing baseline for future fault-tolerant quantum machines, not a present-day quantum CFD product.
The watch item, in the paper's own framing, is whether the same compression survives when the fluid problem stops being Hele-Shaw and the linear systems get bigger and messier. The preprint and the journal version both gesture at broader classes of problems, but the headline benchmark remains the parallel-plate flow. Until the next set of benchmarks lands, this is a working blueprint for hybrid quantum-HPC, not a finished road.