Conversation
Introduce an experimental runner that executes circuits as gate batches over cache-sized state blocks. Each batch selects causally executable raw gates, maps their logical qubits into block-local physical positions, fuses the mapped gates, and executes them across all state blocks in parallel. Bound planning cost by comparing greedy, resident, and upcoming-gate- seeded candidates. Score candidates by circuit progress while accounting for the number of required state-vector swap passes. Apply each set of disjoint qubit-position swaps in one state-vector pass. Keep architecture-specific SIMD lane positions pinned so remapping moves whole chunks on NEON, SSE, AVX2, and AVX-512 layouts. Add QubitMappedState to retain the logical-to-physical qubit layout across execution. The mapped-state API returns without restoring canonical qubit order, while the raw-state compatibility API restores identity before returning. Add a command-line runner for exercising the gate-batch implementation and register the runner and supporting headers with the Bazel build. Register a Makefile target that builds both the gate-batch runner and the existing base runner. Signed-off-by: Xin Xie <xinpin@google.com>
There was a problem hiding this comment.
Code Review
This pull request introduces a gate-batched cache-local simulation runner (QSimGateBatchRunner) along with qubit remapping and layout utilities (QubitLayout, QubitMappedState, and ApplyBitPairSwaps) to optimize simulation performance. It also adds a benchmark application (qsim_gate_batch) and updates the build configurations. Feedback was provided to relax an assertion in run_qsim_gate_batch.h that would otherwise cause failures for small circuits where the number of qubits is less than the SIMD chunk size.
| layout_(layout), | ||
| gate_batch_planner_(partition_.num_state_qubits, | ||
| partition_.block_qubits, chunk_qubits_) { | ||
| assert(partition_.block_qubits >= chunk_qubits_); |
There was a problem hiding this comment.
The assertion assert(partition_.block_qubits >= chunk_qubits_); will fail for small circuits where the number of qubits is less than the SIMD chunk size (e.g., a 2-qubit circuit on an AVX-512 backend where chunk_qubits_ is 4). Since cache blocking is not required for circuits that fit entirely within a single SIMD chunk, we should relax this assertion to allow block_qubits to be smaller than chunk_qubits_ when the total number of state qubits is also smaller than chunk_qubits_.
| assert(partition_.block_qubits >= chunk_qubits_); | |
| assert(partition_.block_qubits >= std::min(chunk_qubits_, partition_.num_state_qubits)); |
|
@xxie24 Thank you for this work. I know it's still a draft PR, but I wanted to mention that a recent failure in the CI workflow has now been fixed. I took the liberty of updating your branch to have the latest changes (I hope that's okay), and now the PR passes all the Ci tests. |
cool. thanks. |
Dynamic scheduling with small chunk sizes caused excessive atomic compare-and-swap synchronization contention on the shared work-sharing counter across high core counts. Switching to schedule(static) eliminates lock contention and preserves sequential hardware stream prefetch performance.
Allow logical CPUs that share a physical core to cooperate on one cache-resident state block. Add the -i/inner_threads option while preserving the existing one-thread-per-block behavior by default. Design decisions: - Split each simulator kernel's iteration range between team lanes with CooperativeFor, and synchronize the lanes after every gate because the next gate consumes the preceding gate's output. - Keep independent-block and SMT-team scheduling explicit, while sharing the block-local gate loop through a compile-time cooperative helper. The non-SMT path therefore pays no barrier or runtime branch overhead. - Discover Linux CPU topology from the process affinity and sysfs package/core identifiers. Order logical CPUs by physical core, pin every team lane explicitly, and reject configurations where the requested siblings are unavailable instead of silently pairing unrelated cores. - Keep non-OpenMP builds working and report a clear error if SMT teams are requested without OpenMP. - Register the cooperative loop and CPU topology headers in the Bazel library. On a Google C4 Intel Xeon Platinum 8581C with 8 physical cores and 16 vCPUs, q30 at d=30, L=18, and f=3 averaged 8.52 seconds with two-way SMT teams. This is about 9% faster than the 8-thread gate-batch run (9.30 seconds) and 2.11x faster than the base runner (17.98 seconds). Team execution reduced the gate phase from 6.32 to 5.93 seconds; remapping averaged 2.59 seconds. Team and non-team q24 runs produced identical amplitudes. Remaining team-SMT work: - Base adaptive block sizing on the number of teams rather than total logical threads. - Separate physical-core remap workers from the larger gate-thread pool. - Add pause/backoff behavior to the per-team spin barrier. - Add topology mocks and OpenMP CI coverage, and extend discovery beyond Linux. - Benchmark more CPU generations, circuit sizes, and affinity configurations. Signed-off-by: Xin Xie <xinpin@google.com>
On multi-node systems, allocating large state vectors on a single NUMA node can create memory controller and channel bottlenecks while remote memory channels remain underutilized. Apply page-level NUMA interleaving (MPOL_INTERLEAVE) via sysfs compute node discovery during state vector allocation on Linux. This balances memory traffic across available memory controllers, noticeably reducing gate execution times on multi-node architectures while gracefully falling back to standard allocation when NUMA support or multiple nodes are unavailable. Tested: bazel test //tests:all
In quantum circuits, phase-shift gates (diagonal unitaries such as Z, S, T, CZ, and Rz) commute with each other. In previous gate-batch planning, a skipped phase gate placed a hard causal barrier on all its qubits, blocking subsequent independent phase gates from packing into the current cache-resident block. Introduce soft barriers for phase gates so that skipped diagonal gates only bar subsequent non-diagonal operations, allowing compatible phase gates to commute past them. Additionally, expose a tunable eviction floor (-e) to enable streaming-optimized spans on wider vector extensions (e.g., floor 5 for AVX-512) and increase lookahead candidate exploration to 64 seeds by default. Tested: bazel test //tests:all
Run: local_script/evaluate_gate_seeds.sh circuits/circuit_q30 12 The harness compares lookahead seeds 0, 1, 8, and 64 with -f 3, -l 19, -e 5, and -x 1. Every run includes the empty and resident candidates. It reports total runtime, batches, executable gates, swaps, planning time, remapping time, and gate time.
The candidate-plan search accumulated a vector of fully materialized
GateBatchPlan objects and then picked the maximum, so peak planner memory
scaled with the number of seeds times the gates admitted by each. Buffer
the lightweight qubit sets instead and evaluate them as they are produced,
keeping only the incumbent best plan alive.
That removes both per-batch heap allocations and makes several helpers
redundant. CollectSeedGateIndices and PlanGateBatchFromQubitSet each had a
single caller and are folded in. GateBatchPlan::operator< became dead when
std::max_element went away.
The eviction floor depends only on constructor inputs, but was recomputed
inside AdditionalRemapSlotsFor, which runs per qubit of every gate of every
candidate seed. Precompute it once; chunk_qubits_ and min_eviction_floor_
were read by nothing else and are dropped as members.
PendingGate exposes `applied` directly rather than through an IsPending /
MarkApplied pair, which also removes a double negative from the three
planner scans.
Selection behavior is unchanged. Ties still resolve to the first candidate
(greedy, then resident, then seeded in circuit order) because the
comparison is strict; the added HasGates() guard fixes an edge case where
the default score of 0.0 would have suppressed a negative-scoring plan.
Net -22 lines. Planning time on q30 at -g 64 improved 7.7% (min of 5,
randomized interleaved A/B), attributable to the hoisted floor computation
and the two removed allocations.
Tested: 32-case equivalence vs parent across {q20,q24} x {-g 0,1,8,64} x
{-x 0,1} x {-l 19,12}, comparing batch/gate/swap counts and amplitudes.
All bit-identical. Built clean at -O3 -fopenmp -march=native -Wall -Wextra.
The file had grown so that its section banners no longer described their
contents. "Gate-batch diagnostics" spanned 433 lines and held 14 logic
members against 4 logging ones, making it the de facto home of remapping,
hot-qubit placement, fusion, and the entire block-execution engine. Nothing
signalled that ExecuteGateBatchOnBlocks lived under a banner saying
diagnostics.
Reorder into four tiers: plain records, then the two driver loops, then
phase implementations in the order the drivers call them, then diagnostics.
SimulateCircuit and ExecuteNextGateBatch are now adjacent, so the algorithm
reads top to bottom instead of being split across 394 lines.
Hoist GateBatchPlanner out of QSimGateBatchRunner. It referenced no
enclosing state and took every input through its constructor, so the
nesting only obscured that it is a self-contained unit; as a namespace-scope
template it is also testable without instantiating the runner. This shrinks
the runner class from 1044 to 795 lines.
Collect all eight Log* members into a trailing section, removing the nine
alternations between logging and logic that previously interrupted the
pipeline. Move IsDiagonalMatrix under the gate-utilities banner it belongs
to, and order SmtTeamBarrier last in the internal namespace, next to the
execution code that uses it.
No behavior change. Every construct moved verbatim; the only content edits
are the template header and <FP> parameterization the planner hoist
requires, plus its alias in the runner.
Tested: verified as a pure move by an indentation-independent line multiset
comparison against the parent, which reports 0 lines removed and exactly 3
added (template <typename FP>, and the two-line planner alias). 32-case
equivalence vs parent across {q20,q24} x {-g 0,1,8,64} x {-x 0,1} x
{-l 19,12}: all batch/gate/swap counts and amplitudes bit-identical. q30
smoke at 32 threads unchanged at 21 batches / 364 gates / 158 swaps. Built
clean at -O3 -fopenmp -march=native -Wall -Wextra.
ExecuteSmtBlockTeams was the longest function in the header at 57 lines,
and its non-OpenMP fallback needed seven (void) casts to silence unused
parameters for a body that could never run.
The fallback casts disappear by guarding the whole function with
#ifdef _OPENMP and moving the error path into its only caller, where
every parameter but team_thread_cpus is still consumed by the
independent-block path.
The team block loop walked a base index and recomputed a per-team offset,
burning empty trailing iterations whenever the base was still in range but
the derived block was not. Striding directly from team_id enumerates the
same block set with no dead iterations and drops the loop-invariant
active check out of the body. This is safe because the barrier inside
ExecuteGatesOnBlock is per-team, so a team's members always iterate
identically and inactive threads belong to wholly inactive teams.
The remaining index arithmetic moves into SmtTeamRole, a pure record with
no OpenMP in it, so the parallel region reads as four steps: assign role,
pin, barrier, run blocks.
Tested: New exhaustive check confirms AssignSmtTeamRole reproduces the
previous inline arithmetic across all 35360 combinations of
inner_threads <= 16, num_threads <= 64, and thread_id, and that active
teams tile blocks 0..n-1 exactly once for every inner_threads <= 8,
num_threads <= 64, num_blocks <= 40.
Tested: 72/72 bit-identical amplitudes and batch/gate/swap counts against
the parent commit over {circuit_q20, circuit_q24} x -g {0,8,64} x
-x {0,1} x (-t/-i) {4/1, 4/2, 8/2, 6/2, 6/3, 12/4}, covering one, two and
three teams and block counts not divisible by the team count.
Tested: 4/4 bit-identical on a build without -fopenmp; -i 2 still reports
that SMT teams require OpenMP.
Tested: Clean build at -O3 -fopenmp -march=native -Wall -Wextra.
Collect the OpenMP-dependent code into one shim block and reject SMT teams without OpenMP during parameter validation, before the state is touched. ExecuteSmtBlockTeams no longer needs an #ifdef.
The planner already computes the eviction floor once; derive the remap slot capacity from it instead of copying it into every candidate set.
Batch remapping and hot-qubit placement now share one helper that moves wanted qubits below a limit, taking victims from the top down. Hot qubits therefore fill the fixed zone in reverse order.
Drop the StartPhaseTimer/AccumulatePhaseSeconds helpers. Timings are always collected (a few calls per batch) and reported at verbosity > 1. Rename the Phase N section banners to Step N, since "phase" has a specific meaning in quantum computing.
They read the runner's state, partition, parameters and executable gates directly instead of taking them as arguments. ExecuteGatesOnBlock takes a block index, and a null barrier replaces the kCooperative template.
One AppendExecutableGates loop with a generic append lambda serves both modes. Qubit order is asserted rather than re-normalized, since raw batch gates and fused gates are already ascending.
qubit_remap.h, run_qsim_gate_batch.h and qsim_gate_batch.cc were missing the license block that every other qsim source file carries.
"Chunk" is not a standard SIMD term. Name the low amplitude-index bits that select a SIMD lane kLaneQubits/lane_qubits (log2 of the lane count), and the [re x N, im x N] unit a lane group. In qubit_remap.h, ChunkBitSwap, PartnerChunk, SwapChunkSpans and floats_per_chunk become GroupBitSwap, PartnerGroup, SwapGroupSpans and floats_per_group. No functional change: output is identical on the Mac, and every x86 state space compiles on WSL.
Drop the unused GateBatchQubitSet alias, replace the lane_qubits_ member with a static constexpr kLaneQubits (it always equals the state space's compile-time constant), and fix the ExecutableGate comment: gates are always fused now.
std::stable_sort keeps equal scores in ascending qubit order, so the comparator no longer needs an explicit tie-breaker.
Its only caller turned the returned list into a mask; return the mask directly and build the candidate list with std::iota.
|
/gemini review |
There was a problem hiding this comment.
Code Review
This pull request introduces a proposal-faithful gate-batched cache-local simulation runner (qsim_gate_batch) along with CPU topology discovery, qubit remapping, and NUMA interleaving support. Feedback on the changes highlights a critical buffer over-read vulnerability in the NUMA interleaving syscall, a performance degradation risk in the SMT barrier spin loop due to a lack of pause/yield instructions, and a missing validation check for num_threads when inner_threads is not active.
| bool ValidateParameters() const { | ||
| if (param_.max_fused_size < 2) { | ||
| IO::errorf("qsim_gate_batch: max_fused_size must be at least 2.\n"); | ||
| return false; | ||
| } | ||
| if (param_.inner_threads <= 1) return true; |
There was a problem hiding this comment.
If param_.inner_threads <= 1, ValidateParameters returns true early without validating param_.num_threads. If param_.num_threads is 0, this can cause undefined behavior or crashes in OpenMP parallel regions (e.g., #pragma omp parallel num_threads(0)). We should validate that param_.num_threads > 0 before the early return.
| bool ValidateParameters() const { | |
| if (param_.max_fused_size < 2) { | |
| IO::errorf("qsim_gate_batch: max_fused_size must be at least 2.\n"); | |
| return false; | |
| } | |
| if (param_.inner_threads <= 1) return true; | |
| bool ValidateParameters() const { | |
| if (param_.max_fused_size < 2) { | |
| IO::errorf("qsim_gate_batch: max_fused_size must be at least 2.\n"); | |
| return false; | |
| } | |
| if (param_.num_threads == 0) { | |
| IO::errorf("qsim_gate_batch: num_threads must be greater than 0.\n"); | |
| return false; | |
| } | |
| if (param_.inner_threads <= 1) return true; |
ValidateParameters returned early for inner_threads <= 1 without checking num_threads, so 0 could reach num_threads(0) in the OpenMP pragmas. Reject it up front, which also covers the SMT-team check. Pass mbind's maxnode as the mask's bit count, as the mbind(2) man page and current libnuma do, instead of the bit count plus one.
Remove casts that the surrounding arithmetic, comparison or parameter type already implies: the swap-loop bound, the plan score, the MatrixShuffle qubit count, the swap counter and the block address. The remaining casts feed printf-style arguments or are otherwise required. Use auto for locals whose initializer already states the type (GetTime timers, element references, shifts of typed constants, arithmetic on typed operands). Counters initialized from literals, loop indices and class members keep their explicit types.
Spell out that both arguments are physical positions, the layout's term for amplitude-index bits.
Rename PhysicalIndexOf to PhysicalAmplitudeIndex, so it no longer reads like PhysicalPositionOf: one maps an amplitude index, the other a qubit. Also remove the unused LogicalToPhysical accessor.
It holds one thread's place in the SMT team grid; AssignSmtTeam builds it. The local at the call site is now assignment instead of role.
The header is mostly QubitLayout; name it after that class.
gate_batch_runner.h holds the planner and pipeline as GateBatchRunner<IO, Fuser, Backend>. A backend owns the state and provides block sizing, swap passes, batch execution and synchronization; its Parameter and defaults (block_qubits, min_eviction_floor) extend the runner's. run_qsim_gate_batch.h keeps the CPU pieces (OpenMP blocks, SMT teams, CooperativeFor) as CpuGateBatchBackend, and QSimGateBatchRunner keeps its Run(param, factory, circuit, state) API. QubitSwap moves to qubit_layout.h with the layout that produces the swaps; qubit_remap.h is only the CPU swap kernel. No behavior change: same batches, swaps and amplitudes.
1164f88 to
87a91d4
Compare
block_qubits becomes tile_qubits, BlockPartition becomes TilePartition, and so on. "Tile" is the usual name for a cache-sized piece of work. The dependency sense of "blocked" (is_qubit_blocked_ and friends) is unchanged. API: Parameter::block_qubits is now tile_qubits. The -l flag is unchanged. No behavior change.
Backends now declare kMaxGateQubits: 6 on the CPU, matching qsim's simulators. The runner rejects a wider raw gate or max_fused_size before it touches the state. max_fused_size alone does not bound raw gates: the fuser passes a wider gate through unchanged.
Both Run overloads sized the layout and backend from circuit.num_qubits. A larger state was then left partly unsimulated, and a smaller one was overrun. They now fail before touching the state when the state (or a mapped state's layout) has a different qubit count.
The adaptive CPU tile size had a floor of max(lane qubits, max_fused_size), so with several threads a small state could get a tile with fewer than two positions above the lanes. A two-qubit gate on non-lane qubits then never fit, and the run failed. The floor is now lane qubits + max_fused_size, or the whole state when that is smaller.
The smallest state is one lane group, so MinSize(0) is 2 * 2^lane_qubits floats. Computing kLaneQubits from it drops the kLaneQubits constant this PR added to the five SIMD state spaces, which are now unchanged from main.
When adaptive tiling lowers tile_qubits below requested_tile_qubits to provide enough tiles for worker threads, the state is already small enough that single-threaded candidate-seed scans can dominate total runtime. Conversely, when tiles reach their full requested size, the state is memory-bandwidth bound and benefits from the full lookahead seed budget to minimize batches and global swap passes. Bound effective lookahead seeds in the tile-clamped regime by both the tile saturation ratio (2^tile_qubits / 2^requested_tile_qubits) and the ratio of per-tile SIMD lane groups to the two gate-scan passes per seed, avoiding hardcoded qubit thresholds while eliminating small-circuit planning overhead. The requested size comes from the new Backend::RequestedTileQubits(): the tile before the CPU backend shrinks it for threads. Tested: - Built and ran tests/run_qsim_test.x and apps/qsim_gate_batch.x - Benchmarked q20..q28 circuits (circuits2, circuits3, heis_q20d20, qft_q20) across T=8,16,32,48,96,192 on Intel C4 (Emerald Rapids) and AMD C4D (Turin Zen 5), verifying up to 2.53x speedup over fixed g=64 on small circuits and zero regression on q26..q28
build_gate_batch.sh
Gate-batched, cache-local simulation on the CPU
Idea
A state-vector simulator is memory-bound. qsim's standard runner makes one full
pass over the state for every fused gate. This PR adds a runner that groups the
circuit into batches of up to K gates and applies each batch to one
cache-sized slice of the state (a tile) at a time, while that slice is in
fast memory. The state is read and written once per batch instead of once per
gate.
K varies from batch to batch. A batch takes as many upcoming gates as the tile
can hold, meaning every qubit they act on fits in the tile. For example,
circuit_q30 runs 1,617 raw gates in 29 batches, about 56 gates per batch.
A tile is 2^L amplitudes, sized for L2 (L = 19 by default). For each batch, the
planner picks gates whose qubits fit in L positions. One swap pass then moves
those qubits into the low L index bits, so every tile is a contiguous slice.
The lane bits are fixed because the swap kernel moves whole SIMD lane groups.
QubitLayoutrecords which logical qubit sits at each physical position, andthe original order is restored at the end.
Parallelism
Tiles are independent within a batch, so each worker takes a whole tile and
applies all K gates to it:
qsim's single-threaded SIMD simulator on it while it stays in that core's L2.
Dynamic scheduling balances fast and slow cores.
once per run. The team shares one tile, so the tile is in L2 once, not twice.
The siblings split each gate's work and meet at a barrier before the next
gate.
Swap passes run over the whole state, not per tile, with OpenMP threads over
contiguous spans.
For small states, the tile shrinks so that every thread still has tiles to
work on, and the planner's lookahead shrinks with it so that planning does not
dominate.
Structure
The pipeline is written against a small backend interface; this PR provides
the CPU backend.
The backend owns the state memory and provides:
The CPU backend reuses qsim's existing AVX-512/AVX2/SSE/NEON/basic simulators
on each tile.
Entry point:
A
QubitMappedStateoverload skips the final restore. App:qsim_gate_batch.cc.Outside the new files,
vectorspace.hnow interleaves state memory across NUMAnodes on Linux.
Results
8-physical-core/16-vCPU
c4-standard-16Intel Xeon Platinum 8581C VM:Limitations
ControlledGateoperations (thecprefix in circuit files), measurementsand channels.