A Jülich Supercomputing Centre preprint compares ten IBM, IQM and Quantinuum processors, testing whether repeated measurement, reset, feed-forward and idling preserve the signals required for quantum error correction.
Quantinuum's trapped-ion processors produced the clearest signal in a cross-platform test of the operations that quantum error correction depends on. The benchmark preserved algorithmic signal on a 91-data-qubit triangular color code while requiring 45 mid-circuit measurements per layer across 98 physical qubits.
Researchers at Germany's Jülich Supercomputing Centre presented the work in the preprint Jülich arXiv preprint, first posted on 5 October 2026 and updated on 6 October. The study compares ten commercial quantum processing units from Quantinuum, IBM and IQM, rather than evaluating only a single hardware family. It remains a preprint, not a peer-reviewed journal publication.
The benchmark examines four operations that sit between an ideal error-correcting code and a working processor: mid-circuit measurement, qubit reset, dynamic feed-forward control and spectator idling. These operations are not secondary housekeeping. A QEC cycle must measure error syndromes without ending the computation, reset qubits for reuse, condition later operations on earlier measurement outcomes and prevent supposedly idle qubits from accumulating damaging errors.
A processor can therefore look strong on isolated gate metrics while losing useful signal when these controls are repeated across a wide and deep circuit. This systems-level perspective is closer to the operational questions now emphasized across quantum-computing research, including work associated with institutions such as MIT and discussions in journals such as Nature, where device performance is increasingly assessed through complete workloads rather than isolated component specifications.
Rather than relying on large numbers of samples, the researchers used a low-sampling algorithmic benchmark built around a Linear-Ramp Quantum Approximate Optimization Algorithm framework. The LR-QAOA construction used fixed linear parameters, embedded the interaction and check structures of distance-scaled surface codes, triangular color codes and bivariate-bicycle quantum low-density parity-check codes, and converted the decline in algorithmic performance with increasing depth into an effective two-qubit depolarizing error.
The measured quantity is therefore an algorithm-level view of how control operations combine under pressure, not a catalogue of one-qubit or two-qubit specifications. It also probes a workload shaped by QEC circuits and should not be read as a universal ranking of every capability of the processors.
The experiments reached circuits containing as many as 2,950 mid-circuit measurement operations. On Quantinuum's Helios-1 and H2-1 systems, the researchers tested surface-code structures with up to 81 data qubits, triangular color-code structures with up to 91 qubits and bivariate-bicycle qLDPC structures with up to 48 qubits. Some individual runs used as many as 480 mid-circuit measurements.
For the 91-qubit triangular color-code test, the reported algorithmic-signal retention was r = 0.733. The circuit required 45 mid-circuit measurements per layer and used 98 physical qubits. In the comparison, the effective contribution of mid-circuit measurement on Quantinuum systems was estimated at approximately 0.22-0.49%, a level described as comparable to two-qubit-operation errors. The corresponding range reported for the tested IBM Heron and Eagle processors was about 1.1-3.2%, roughly an order of magnitude higher than the Quantinuum baseline.
Those figures do not establish a universal hardware hierarchy. They describe the behavior of particular systems under the selected circuits, compilation choices and measurement protocol. Architecture matters: trapped-ion processors and superconducting processors expose different connectivity, timing, reset and control constraints, so the same code family can impose different practical burdens on each platform.
The researchers also tested whether the low-shot metric tracked a more direct memory result. On IBM's ibm_phoenix, they compared the LR-QAOA result with an 11-patch distance-three surface-code memory experiment and found a Spearman correlation of ρs = 0.83-0.88 between the benchmark measure and observed logical-state survival probabilities.
That comparison is important because a compact benchmark is useful only if its score corresponds to behavior that matters for encoded information. The reported relationship connects the algorithmic signal to logical memory survival on one IBM processor. It does not turn the benchmark into a logical error rate for every device or establish that the tested systems operated as fault-tolerant quantum computers.
The distinction also separates hardware-level error correction from the cryptographic migration discussed in earlier coverage. Post-quantum cryptography runs on conventional computing infrastructure, whereas this study asks whether physical quantum hardware can execute the repeated measurements and conditional controls needed to protect quantum information.
The study's strongest contribution is methodological: it offers a relatively low-cost preliminary check of QEC suitability before a laboratory undertakes a full logical-memory experiment. Its results indicate that the LR-QAOA metric can track a memory-related outcome in the reported IBM experiment and distinguish performance among the tested platforms. The supplied findings emphasize correlation and effective error estimates; they do not provide a basis here for adding unreported p-values or confidence intervals.
But the experiment is still a benchmark of QEC building blocks. The work does not report a complete logical processor, a growing logical-qubit advantage, real-time decoder performance, a fault-tolerant logical gate or a useful error-corrected algorithm. Nor does it establish independent replication beyond the processors and experiments included in the study.
That boundary matters. Surface codes, color codes and qLDPC codes impose different interaction and check structures, so performance depends on how each processor handles measurement, reset, conditional control and idling under the chosen circuit. The benchmark can reveal accumulated algorithmic damage, but it cannot by itself settle which architecture will deliver the lowest overhead or the most manufacturable fault-tolerant system.
The evidence nevertheless points to a practical criterion for quantum hardware: raw qubit count is not enough. A processor must preserve usable information while repeatedly measuring, resetting and conditionally controlling many qubits. This Jülich study tests that operational burden directly, offering a sharper screening tool for future error-correction experiments rather than a claim that fault tolerance has already been achieved.
Quantum error correction encodes one logical qubit across multiple physical qubits so that patterns of errors can be detected and potentially corrected. A physical-qubit benchmark such as this one tests whether the hardware can perform the required syndrome-extraction operations reliably; it does not automatically demonstrate a protected logical qubit. Evidence for fault tolerance would require logical performance to improve as encoding resources grow, together with the control, decoding and repeated-cycle results needed to show that errors are being reduced rather than merely measured.