Add zero-copy GPU state vector access (#836) - #1100
Jonathan-tWang wants to merge 9 commits into
Conversation
This PR adds an opt-in zero-copy API, `QSimSimulator.simulate_into_device_array(...)`, to extract the final state vector from GPU simulation without copying data from GPU to host memory (issue quantumlib#836). The returned `DeviceStateVector` object owns the GPU allocation and exposes it through the CUDA Array Interface (`__cuda_array_interface__`, v3), allowing downstream GPU frameworks (CuPy, PyTorch, Numba) to consume the device buffer directly without a device -> host -> device round trip.
|
Thanks for your pull request! It looks like this may be your first contribution to a Google open source project. Before we can look at your pull request, you'll need to sign a Contributor License Agreement (CLA). View this failed invocation of the CLA check for more information. For the most up to date status, view the checks section at the bottom of the pull request. |
There was a problem hiding this comment.
Code Review
This pull request implements zero-copy device state-vector bindings, allowing GPU-based simulations to retain their final state in device memory and expose it via the CUDA Array Interface for direct consumption by libraries like CuPy. Feedback on these changes highlights a potential out-of-bounds read vulnerability when passing non-contiguous NumPy arrays as initial states, which can be resolved by ensuring C-contiguity. Additionally, it is recommended to use std::make_unique instead of direct new expressions when instantiating std::unique_ptrs to comply with the Google C++ Style Guide.
- Ensure initial_state NumPy array is C-contiguous (via np.ascontiguousarray) before creating float32 view to prevent out-of-bounds reads on non-contiguous array slices. - Use std::make_unique<DeviceStateVector> instead of bare new expression per Google C++ Style Guide. - Add test_cirq_qsim_gpu_simulate_into_device_array_with_non_contiguous_input_state test case.
f8f86e3 to
ad041bf
Compare
ad041bf to
9d255a8
Compare
The *_avx512_test targets are compiled with -march=native, so their object code depends on the build machine's CPU while their compile command line does not. Machines with the same toolchain therefore compute the same Bazel action key for these compiles, and a shared disk cache (CI restores ~/.cache/bazel/disk_cache across jobs) can hand an object built for an AVX-512 CPU to a runner without AVX-512, where the test dies with SIGILL. compiler_probe now also records NATIVE_ARCH_ID, a hash of the compiler's predefined macros under -march=native (which name the instruction set extensions it may use), and tests/BUILD adds it to avx512_copts as -DQSIM_NATIVE_ARCH_ID=<id>. Machines whose -march=native differs now get different action keys for these four compiles, so each one builds its own objects. The macro is unused, so the generated code does not change.
The destructor released the states by assigning StateSpace::Null() to them. For cuStateVecEx that frees nothing: the Vector's move assignment (lib/vectorspace_custatevecex.h) overwrites its descriptor without destroying it. So with gpu_mode=2, every call that goes through SimulatorHelper leaked its state vector and the device memory it owns: simulate(), expectation values, run() when it samples the final state, and simulate_into_device_array() + free(). Move the states into a temporary instead, so that their destructors run while the device guard (in nvcc builds) keeps the owning device current. Add a GPU test that repeated simulations do not lose device memory, for each GPU mode and for both the host and the device-array paths. It needs a GPU build and cupy, and it skips under pytest-xdist with more than one worker because it measures device-wide free memory.
This reverts commit 5f552a1. The change is unrelated to issue quantumlib#836, and the CI failure it was meant to prevent was never traced to its cause. Keep this PR focused on quantumlib#836.
This PR adds an opt-in zero-copy API,
QSimSimulator.simulate_into_device_array(...), to extract the final state vector from GPU simulation without copying data from GPU to host memory (issue #836).The returned
DeviceStateVectorobject owns the GPU allocation and exposes it through the CUDA Array Interface (__cuda_array_interface__, v3), allowing downstream GPU frameworks (CuPy, PyTorch, Numba) to consume the device buffer directly without a device -> host -> device round trip.