Skip to content

Add gate-batched cache-local runner across SIMD backends - #1092

Draft
xxie24 wants to merge 55 commits into
quantumlib:mainfrom
xxie24:blocked-v2
Draft

xxie24 wants to merge 55 commits into
quantumlib:mainfrom
xxie24:blocked-v2

Conversation

@xxie24

@xxie24 xxie24 commented Jul 23, 2026 •

Copy link
Copy Markdown
Contributor

build_gate_batch.sh

Gate-batched, cache-local simulation on the CPU

Idea

A state-vector simulator is memory-bound. qsim's standard runner makes one full
pass over the state for every fused gate. This PR adds a runner that groups the
circuit into batches of up to K gates and applies each batch to one
cache-sized slice of the state (a tile) at a time, while that slice is in
fast memory. The state is read and written once per batch instead of once per
gate.

K varies from batch to batch. A batch takes as many upcoming gates as the tile
can hold, meaning every qubit they act on fits in the tile. For example,
circuit_q30 runs 1,617 raw gates in 29 batches, about 56 gates per batch.

standard:    gate 1 → full pass │ gate 2 → full pass │ ... │ gate g → full pass
gate batch:  [swap pass] + for each tile: load → gates 1..K → store   (per batch)

A tile is 2^L amplitudes, sized for L2 (L = 19 by default). For each batch, the
planner picks gates whose qubits fit in L positions. One swap pass then moves
those qubits into the low L index bits, so every tile is a contiguous slice.

amplitude index bits:   n-1 ........ L │ L-1 ........ lane │ lane-1 ... 0
                        outside tile   │ tile (remappable) │ SIMD lanes

The lane bits are fixed because the swap kernel moves whole SIMD lane groups.
QubitLayout records which logical qubit sits at each physical position, and
the original order is restored at the end.

Parallelism

Tiles are independent within a batch, so each worker takes a whole tile and
applies all K gates to it:

batch b:   tile 0 │ tile 1 │ tile 2 │ ... │ tile 2^(n-L)-1
              ↓        ↓        ↓               ↓
           thread   thread   thread   ...    (OpenMP, dynamic scheduling)
  or:      team     team     team     ...    (two SMT siblings split each gate)
  • Independent threads: each OpenMP thread takes one tile at a time and runs
    qsim's single-threaded SIMD simulator on it while it stays in that core's L2.
    Dynamic scheduling balances fast and slow cores.
  • SMT teams (Linux): the two hardware threads of a core form a team, pinned
    once per run. The team shares one tile, so the tile is in L2 once, not twice.
    The siblings split each gate's work and meet at a barrier before the next
    gate.

Swap passes run over the whole state, not per tile, with OpenMP threads over
contiguous spans.

For small states, the tile shrinks so that every thread still has tiles to
work on, and the planner's lookahead shrinks with it so that planning does not
dominate.

Structure

The pipeline is written against a small backend interface; this PR provides
the CPU backend.

gate_batch_runner.h        GateBatchRunner<IO, Fuser, Backend>
  planner, remap bookkeeping, per-batch fusion, restore
  qubit_layout.h           QubitLayout, QubitMappedState, QubitSwap
             │
             │  Backend interface
             │
run_qsim_gate_batch.h      CpuGateBatchBackend, QSimGateBatchRunner
 ├ qubit_remap.h           swap passes
 ├ qsim SIMD simulator     applied to each tile
 └ cpu_thread_topology.h   SMT sibling teams

The backend owns the state memory and provides:

struct Parameter;                  // merged into the runner's Parameter
kDefaultTileQubits, kDefaultEvictionFloor, kMaxGateQubits
TileQubits(), RequestedTileQubits(), LaneQubits()
Prepare()                          // validate options, set up threads
ApplySwaps(swaps)                  // one pass of disjoint bit-pair swaps
ExecuteBatch(gates)                // every gate on every tile
Synchronize()                      // wait for queued work

The CPU backend reuses qsim's existing AVX-512/AVX2/SSE/NEON/basic simulators
on each tile.

Entry point:

QSimGateBatchRunner<IO, Fuser, Factory>::Run(param, factory, circuit, state)

A QubitMappedState overload skips the final restore. App:
qsim_gate_batch.cc.

Outside the new files, vectorspace.h now interleaves state memory across NUMA
nodes on Linux.

Results

8-physical-core/16-vCPU c4-standard-16 Intel Xeon Platinum 8581C VM:

Configuration CPU allocation Total (s) Gate application (s) Remap (s) Speedup vs 8-core base
Base runner 8 physical cores 17.980 — — 1.00x
Gate-batch, physical only 8 independent workers 9.297 6.319 2.976 1.93x
Gate-batch, ordinary SMT 16 independent workers 9.257 6.762 2.493 1.94x
Gate-batch, team SMT 8 two-sibling teams 8.523 5.930 2.592 2.11x

Limitations

  • Supports any matrix gate (CNOT, CZ, SWAP, fSim, …). Not yet supported:
    ControlledGate operations (the c prefix in circuit files), measurements
    and channels.
  • The correctness tests are not yet registered with Make/Bazel (follow-up).

Introduce an experimental runner that executes circuits as gate batches
over cache-sized state blocks. Each batch selects causally executable
raw gates, maps their logical qubits into block-local physical positions,
fuses the mapped gates, and executes them across all state blocks in
parallel.

Bound planning cost by comparing greedy, resident, and upcoming-gate-
seeded candidates. Score candidates by circuit progress while accounting
for the number of required state-vector swap passes.

Apply each set of disjoint qubit-position swaps in one state-vector pass.
Keep architecture-specific SIMD lane positions pinned so remapping moves
whole chunks on NEON, SSE, AVX2, and AVX-512 layouts.

Add QubitMappedState to retain the logical-to-physical qubit layout
across execution. The mapped-state API returns without restoring
canonical qubit order, while the raw-state compatibility API restores
identity before returning.

Add a command-line runner for exercising the gate-batch implementation
and register the runner and supporting headers with the Bazel build.
Register a Makefile target that builds both the gate-batch runner and
the existing base runner.

Signed-off-by: Xin Xie <xinpin@google.com>
@github-actions github-actions Bot added the size: XL lines changed >1000 label Jul 23, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a gate-batched cache-local simulation runner (QSimGateBatchRunner) along with qubit remapping and layout utilities (QubitLayout, QubitMappedState, and ApplyBitPairSwaps) to optimize simulation performance. It also adds a benchmark application (qsim_gate_batch) and updates the build configurations. Feedback was provided to relax an assertion in run_qsim_gate_batch.h that would otherwise cause failures for small circuits where the number of qubits is less than the SIMD chunk size.

Comment thread lib/run_qsim_gate_batch.h Outdated
layout_(layout),
gate_batch_planner_(partition_.num_state_qubits,
partition_.block_qubits, chunk_qubits_) {
assert(partition_.block_qubits >= chunk_qubits_);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The assertion assert(partition_.block_qubits >= chunk_qubits_); will fail for small circuits where the number of qubits is less than the SIMD chunk size (e.g., a 2-qubit circuit on an AVX-512 backend where chunk_qubits_ is 4). Since cache blocking is not required for circuits that fit entirely within a single SIMD chunk, we should relax this assertion to allow block_qubits to be smaller than chunk_qubits_ when the total number of state qubits is also smaller than chunk_qubits_.

Suggested change
assert(partition_.block_qubits >= chunk_qubits_);
assert(partition_.block_qubits >= std::min(chunk_qubits_, partition_.num_state_qubits));

@xxie24
xxie24 marked this pull request as draft July 23, 2026 21:12
@mhucka

mhucka commented Aug 15, 2026

Copy link
Copy Markdown
Collaborator

@xxie24 Thank you for this work. I know it's still a draft PR, but I wanted to mention that a recent failure in the CI workflow has now been fixed. I took the liberty of updating your branch to have the latest changes (I hope that's okay), and now the PR passes all the Ci tests.

@xxie24

xxie24 commented Aug 17, 2026

Copy link
Copy Markdown
Contributor Author

@xxie24 Thank you for this work. I know it's still a draft PR, but I wanted to mention that a recent failure in the CI workflow has now been fixed. I took the liberty of updating your branch to have the latest changes (I hope that's okay), and now the PR passes all the Ci tests.

cool. thanks.

Xin Xie added 4 commits August 16, 2026 21:57
Dynamic scheduling with small chunk sizes caused excessive atomic compare-and-swap
synchronization contention on the shared work-sharing counter across high core counts.
Switching to schedule(static) eliminates lock contention and preserves sequential
hardware stream prefetch performance.
Xin Xie added 2 commits August 24, 2026 07:15
Allow logical CPUs that share a physical core to cooperate on one cache-resident state block. Add the -i/inner_threads option while preserving the existing one-thread-per-block behavior by default.

Design decisions:

- Split each simulator kernel's iteration range between team lanes with CooperativeFor, and synchronize the lanes after every gate because the next gate consumes the preceding gate's output.
- Keep independent-block and SMT-team scheduling explicit, while sharing the block-local gate loop through a compile-time cooperative helper. The non-SMT path therefore pays no barrier or runtime branch overhead.
- Discover Linux CPU topology from the process affinity and sysfs package/core identifiers. Order logical CPUs by physical core, pin every team lane explicitly, and reject configurations where the requested siblings are unavailable instead of silently pairing unrelated cores.
- Keep non-OpenMP builds working and report a clear error if SMT teams are requested without OpenMP.
- Register the cooperative loop and CPU topology headers in the Bazel library.

On a Google C4 Intel Xeon Platinum 8581C with 8 physical cores and 16 vCPUs, q30 at d=30, L=18, and f=3 averaged 8.52 seconds with two-way SMT teams. This is about 9% faster than the 8-thread gate-batch run (9.30 seconds) and 2.11x faster than the base runner (17.98 seconds). Team execution reduced the gate phase from 6.32 to 5.93 seconds; remapping averaged 2.59 seconds. Team and non-team q24 runs produced identical amplitudes.

Remaining team-SMT work:

- Base adaptive block sizing on the number of teams rather than total logical threads.
- Separate physical-core remap workers from the larger gate-thread pool.
- Add pause/backoff behavior to the per-team spin barrier.
- Add topology mocks and OpenMP CI coverage, and extend discovery beyond Linux.
- Benchmark more CPU generations, circuit sizes, and affinity configurations.

Signed-off-by: Xin Xie <xinpin@google.com>
On multi-node systems, allocating large state vectors on a single NUMA
node can create memory controller and channel bottlenecks while remote
memory channels remain underutilized.

Apply page-level NUMA interleaving (MPOL_INTERLEAVE) via sysfs compute
node discovery during state vector allocation on Linux. This balances
memory traffic across available memory controllers, noticeably reducing
gate execution times on multi-node architectures while gracefully falling
back to standard allocation when NUMA support or multiple nodes are
unavailable.

Tested: bazel test //tests:all
xinpin added 15 commits September 15, 2026 15:00
In quantum circuits, phase-shift gates (diagonal unitaries such as Z,
S, T, CZ, and Rz) commute with each other. In previous gate-batch
planning, a skipped phase gate placed a hard causal barrier on all its
qubits, blocking subsequent independent phase gates from packing into the
current cache-resident block.

Introduce soft barriers for phase gates so that skipped diagonal gates
only bar subsequent non-diagonal operations, allowing compatible phase
gates to commute past them. Additionally, expose a tunable eviction
floor (-e) to enable streaming-optimized spans on wider vector extensions
(e.g., floor 5 for AVX-512) and increase lookahead candidate exploration
to 64 seeds by default.

Tested: bazel test //tests:all
Run: local_script/evaluate_gate_seeds.sh circuits/circuit_q30 12

The harness compares lookahead seeds 0, 1, 8, and 64 with -f 3, -l 19, -e 5, and -x 1. Every run includes the empty and resident candidates. It reports total runtime, batches, executable gates, swaps, planning time, remapping time, and gate time.
The candidate-plan search accumulated a vector of fully materialized
GateBatchPlan objects and then picked the maximum, so peak planner memory
scaled with the number of seeds times the gates admitted by each. Buffer
the lightweight qubit sets instead and evaluate them as they are produced,
keeping only the incumbent best plan alive.

That removes both per-batch heap allocations and makes several helpers
redundant. CollectSeedGateIndices and PlanGateBatchFromQubitSet each had a
single caller and are folded in. GateBatchPlan::operator< became dead when
std::max_element went away.

The eviction floor depends only on constructor inputs, but was recomputed
inside AdditionalRemapSlotsFor, which runs per qubit of every gate of every
candidate seed. Precompute it once; chunk_qubits_ and min_eviction_floor_
were read by nothing else and are dropped as members.

PendingGate exposes `applied` directly rather than through an IsPending /
MarkApplied pair, which also removes a double negative from the three
planner scans.

Selection behavior is unchanged. Ties still resolve to the first candidate
(greedy, then resident, then seeded in circuit order) because the
comparison is strict; the added HasGates() guard fixes an edge case where
the default score of 0.0 would have suppressed a negative-scoring plan.

Net -22 lines. Planning time on q30 at -g 64 improved 7.7% (min of 5,
randomized interleaved A/B), attributable to the hoisted floor computation
and the two removed allocations.

Tested: 32-case equivalence vs parent across {q20,q24} x {-g 0,1,8,64} x
{-x 0,1} x {-l 19,12}, comparing batch/gate/swap counts and amplitudes.
All bit-identical. Built clean at -O3 -fopenmp -march=native -Wall -Wextra.
The file had grown so that its section banners no longer described their
contents. "Gate-batch diagnostics" spanned 433 lines and held 14 logic
members against 4 logging ones, making it the de facto home of remapping,
hot-qubit placement, fusion, and the entire block-execution engine. Nothing
signalled that ExecuteGateBatchOnBlocks lived under a banner saying
diagnostics.

Reorder into four tiers: plain records, then the two driver loops, then
phase implementations in the order the drivers call them, then diagnostics.
SimulateCircuit and ExecuteNextGateBatch are now adjacent, so the algorithm
reads top to bottom instead of being split across 394 lines.

Hoist GateBatchPlanner out of QSimGateBatchRunner. It referenced no
enclosing state and took every input through its constructor, so the
nesting only obscured that it is a self-contained unit; as a namespace-scope
template it is also testable without instantiating the runner. This shrinks
the runner class from 1044 to 795 lines.

Collect all eight Log* members into a trailing section, removing the nine
alternations between logging and logic that previously interrupted the
pipeline. Move IsDiagonalMatrix under the gate-utilities banner it belongs
to, and order SmtTeamBarrier last in the internal namespace, next to the
execution code that uses it.

No behavior change. Every construct moved verbatim; the only content edits
are the template header and <FP> parameterization the planner hoist
requires, plus its alias in the runner.

Tested: verified as a pure move by an indentation-independent line multiset
comparison against the parent, which reports 0 lines removed and exactly 3
added (template <typename FP>, and the two-line planner alias). 32-case
equivalence vs parent across {q20,q24} x {-g 0,1,8,64} x {-x 0,1} x
{-l 19,12}: all batch/gate/swap counts and amplitudes bit-identical. q30
smoke at 32 threads unchanged at 21 batches / 364 gates / 158 swaps. Built
clean at -O3 -fopenmp -march=native -Wall -Wextra.
ExecuteSmtBlockTeams was the longest function in the header at 57 lines,
and its non-OpenMP fallback needed seven (void) casts to silence unused
parameters for a body that could never run.

The fallback casts disappear by guarding the whole function with
#ifdef _OPENMP and moving the error path into its only caller, where
every parameter but team_thread_cpus is still consumed by the
independent-block path.

The team block loop walked a base index and recomputed a per-team offset,
burning empty trailing iterations whenever the base was still in range but
the derived block was not. Striding directly from team_id enumerates the
same block set with no dead iterations and drops the loop-invariant
active check out of the body. This is safe because the barrier inside
ExecuteGatesOnBlock is per-team, so a team's members always iterate
identically and inactive threads belong to wholly inactive teams.

The remaining index arithmetic moves into SmtTeamRole, a pure record with
no OpenMP in it, so the parallel region reads as four steps: assign role,
pin, barrier, run blocks.

Tested: New exhaustive check confirms AssignSmtTeamRole reproduces the
previous inline arithmetic across all 35360 combinations of
inner_threads <= 16, num_threads <= 64, and thread_id, and that active
teams tile blocks 0..n-1 exactly once for every inner_threads <= 8,
num_threads <= 64, num_blocks <= 40.
Tested: 72/72 bit-identical amplitudes and batch/gate/swap counts against
the parent commit over {circuit_q20, circuit_q24} x -g {0,8,64} x
-x {0,1} x (-t/-i) {4/1, 4/2, 8/2, 6/2, 6/3, 12/4}, covering one, two and
three teams and block counts not divisible by the team count.
Tested: 4/4 bit-identical on a build without -fopenmp; -i 2 still reports
that SMT teams require OpenMP.
Tested: Clean build at -O3 -fopenmp -march=native -Wall -Wextra.
Collect the OpenMP-dependent code into one shim block and reject SMT
teams without OpenMP during parameter validation, before the state is
touched. ExecuteSmtBlockTeams no longer needs an #ifdef.
The planner already computes the eviction floor once; derive the remap
slot capacity from it instead of copying it into every candidate set.
Batch remapping and hot-qubit placement now share one helper that moves
wanted qubits below a limit, taking victims from the top down. Hot
qubits therefore fill the fixed zone in reverse order.
Drop the StartPhaseTimer/AccumulatePhaseSeconds helpers. Timings are
always collected (a few calls per batch) and reported at verbosity > 1.
Rename the Phase N section banners to Step N, since "phase" has a
specific meaning in quantum computing.
They read the runner's state, partition, parameters and executable gates
directly instead of taking them as arguments. ExecuteGatesOnBlock takes a
block index, and a null barrier replaces the kCooperative template.
One AppendExecutableGates loop with a generic append lambda serves both
modes. Qubit order is asserted rather than re-normalized, since raw
batch gates and fused gates are already ascending.
Xin Xie added 3 commits September 26, 2026 22:08
qubit_remap.h, run_qsim_gate_batch.h and qsim_gate_batch.cc were missing
the license block that every other qsim source file carries.
"Chunk" is not a standard SIMD term. Name the low amplitude-index bits
that select a SIMD lane kLaneQubits/lane_qubits (log2 of the lane
count), and the [re x N, im x N] unit a lane group. In qubit_remap.h,
ChunkBitSwap, PartnerChunk, SwapChunkSpans and floats_per_chunk become
GroupBitSwap, PartnerGroup, SwapGroupSpans and floats_per_group.

No functional change: output is identical on the Mac, and every x86
state space compiles on WSL.
Drop the unused GateBatchQubitSet alias, replace the lane_qubits_ member
with a static constexpr kLaneQubits (it always equals the state space's
compile-time constant), and fix the ExecutableGate comment: gates are
always fused now.
Xin Xie added 3 commits September 26, 2026 22:30
std::stable_sort keeps equal scores in ascending qubit order, so the
comparator no longer needs an explicit tie-breaker.
Its only caller turned the returned list into a mask; return the mask
directly and build the candidate list with std::iota.
@xxie24

xxie24 commented Sep 27, 2026

Copy link
Copy Markdown
Contributor Author

/gemini review

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a proposal-faithful gate-batched cache-local simulation runner (qsim_gate_batch) along with CPU topology discovery, qubit remapping, and NUMA interleaving support. Feedback on the changes highlights a critical buffer over-read vulnerability in the NUMA interleaving syscall, a performance degradation risk in the SMT barrier spin loop due to a lack of pause/yield instructions, and a missing validation check for num_threads when inner_threads is not active.

Comment thread lib/vectorspace.h
Comment thread lib/run_qsim_gate_batch.h
Comment thread lib/run_qsim_gate_batch.h Outdated
Comment on lines +1054 to +1059
bool ValidateParameters() const {
if (param_.max_fused_size < 2) {
IO::errorf("qsim_gate_batch: max_fused_size must be at least 2.\n");
return false;
}
if (param_.inner_threads <= 1) return true;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

If param_.inner_threads <= 1, ValidateParameters returns true early without validating param_.num_threads. If param_.num_threads is 0, this can cause undefined behavior or crashes in OpenMP parallel regions (e.g., #pragma omp parallel num_threads(0)). We should validate that param_.num_threads > 0 before the early return.

Suggested change
bool ValidateParameters() const {
if (param_.max_fused_size < 2) {
IO::errorf("qsim_gate_batch: max_fused_size must be at least 2.\n");
return false;
}
if (param_.inner_threads <= 1) return true;
bool ValidateParameters() const {
if (param_.max_fused_size < 2) {
IO::errorf("qsim_gate_batch: max_fused_size must be at least 2.\n");
return false;
}
if (param_.num_threads == 0) {
IO::errorf("qsim_gate_batch: num_threads must be greater than 0.\n");
return false;
}
if (param_.inner_threads <= 1) return true;

Xin Xie added 8 commits September 27, 2026 09:08
ValidateParameters returned early for inner_threads <= 1 without checking
num_threads, so 0 could reach num_threads(0) in the OpenMP pragmas.
Reject it up front, which also covers the SMT-team check.

Pass mbind's maxnode as the mask's bit count, as the mbind(2) man page
and current libnuma do, instead of the bit count plus one.
Remove casts that the surrounding arithmetic, comparison or parameter
type already implies: the swap-loop bound, the plan score, the
MatrixShuffle qubit count, the swap counter and the block address. The
remaining casts feed printf-style arguments or are otherwise required.

Use auto for locals whose initializer already states the type (GetTime
timers, element references, shifts of typed constants, arithmetic on
typed operands). Counters initialized from literals, loop indices and
class members keep their explicit types.
Spell out that both arguments are physical positions, the layout's term
for amplitude-index bits.
Rename PhysicalIndexOf to PhysicalAmplitudeIndex, so it no longer reads
like PhysicalPositionOf: one maps an amplitude index, the other a qubit.
Also remove the unused LogicalToPhysical accessor.
It holds one thread's place in the SMT team grid; AssignSmtTeam builds
it. The local at the call site is now assignment instead of role.
The header is mostly QubitLayout; name it after that class.
gate_batch_runner.h holds the planner and pipeline as
GateBatchRunner<IO, Fuser, Backend>. A backend owns the state and provides
block sizing, swap passes, batch execution and synchronization; its
Parameter and defaults (block_qubits, min_eviction_floor) extend the
runner's.

run_qsim_gate_batch.h keeps the CPU pieces (OpenMP blocks, SMT teams,
CooperativeFor) as CpuGateBatchBackend, and QSimGateBatchRunner keeps its
Run(param, factory, circuit, state) API. QubitSwap moves to qubit_layout.h
with the layout that produces the swaps; qubit_remap.h is only the CPU
swap kernel.

No behavior change: same batches, swaps and amplitudes.
@xxie24
xxie24 force-pushed the blocked-v2 branch 6 times, most recently from 1164f88 to 87a91d4 Compare September 29, 2026 05:35
Xin Xie added 6 commits September 29, 2026 20:01
block_qubits becomes tile_qubits, BlockPartition becomes TilePartition,
and so on. "Tile" is the usual name for a cache-sized piece of work. The
dependency sense of
"blocked" (is_qubit_blocked_ and friends) is unchanged.

API: Parameter::block_qubits is now tile_qubits. The -l flag is
unchanged. No behavior change.
Backends now declare kMaxGateQubits: 6 on the CPU, matching qsim's
simulators. The runner rejects a wider raw gate or max_fused_size before
it touches the state. max_fused_size alone does not bound raw gates: the
fuser passes a wider gate through unchanged.
Both Run overloads sized the layout and backend from circuit.num_qubits.
A larger state was then left partly unsimulated, and a smaller one was
overrun. They now fail before touching the state when the state (or a
mapped state's layout) has a different qubit count.
The adaptive CPU tile size had a floor of max(lane qubits, max_fused_size),
so with several threads a small state could get a tile with fewer than two
positions above the lanes. A two-qubit gate on non-lane qubits then never
fit, and the run failed. The floor is now lane qubits + max_fused_size,
or the whole state when that is smaller.
The smallest state is one lane group, so MinSize(0) is 2 * 2^lane_qubits
floats. Computing kLaneQubits from it drops the kLaneQubits constant this
PR added to the five SIMD state spaces, which are now unchanged from main.
When adaptive tiling lowers tile_qubits below requested_tile_qubits to
provide enough tiles for worker threads, the state is already small
enough that single-threaded candidate-seed scans can dominate total
runtime. Conversely, when tiles reach their full requested size, the
state is memory-bandwidth bound and benefits from the full lookahead
seed budget to minimize batches and global swap passes.

Bound effective lookahead seeds in the tile-clamped regime by both the
tile saturation ratio (2^tile_qubits / 2^requested_tile_qubits) and the
ratio of per-tile SIMD lane groups to the two gate-scan passes per seed,
avoiding hardcoded qubit thresholds while eliminating small-circuit
planning overhead.

The requested size comes from the new Backend::RequestedTileQubits(): the
tile before the CPU backend shrinks it for threads.

Tested:
- Built and ran tests/run_qsim_test.x and apps/qsim_gate_batch.x
- Benchmarked q20..q28 circuits (circuits2, circuits3, heis_q20d20,
  qft_q20) across T=8,16,32,48,96,192 on Intel C4 (Emerald Rapids) and
  AMD C4D (Turin Zen 5), verifying up to 2.53x speedup over fixed g=64
  on small circuits and zero regression on q26..q28
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size: XL lines changed >1000

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants