Repository navigation
Provide a way to avoid copying state vector from Cirq #893
Description
Activity
- addedcontributors welcomeHelp with this would be appreciatedHelp with this would be appreciateda: qsimcirqInvolves the qsim-Cirq integration layerInvolves the qsim-Cirq integration layera: pythonInvolves the Python code in qsimInvolves the Python code in qsima: performanceInvolves performance problems or improvementsInvolves performance problems or improvementsP2Priority: medium priorityPriority: medium priority
on Sep 23, 2025 - addedno QC knowledge neededDoes not require knowledge of quantum computingDoes not require knowledge of quantum computing
on Nov 3, 2025 Hi @mhucka . In
cirq-core/cirq/sim/state_vector_simulation_state.py.__BufferedStateVector.create(), the following condition evaluates truthy, which forces the copy of the array to be made, causing the above error:if np.may_share_memory(state_vector, initial_state): state_vector = state_vector.copy()It appears to me that this is to protect the integrity of the original
initial_stateargument passed to_BufferedStateVector.create()against unwanted mutations.Could we create a flag in
cirq-core/cirq/sim/state_vector_simulation_state.py.StateVectorSimulationStatethat allows the caller to choose whether they care about the integrity of theinitial_statearray or not. If they don't care, then skip the copy condition.Something like:
In StateVectorSimulationState:
def __init__( self, *, available_buffer: np.ndarray | None = None, prng: np.random.RandomState | None = None, qubits: Sequence[cirq.Qid] | None = None, initial_state: np.ndarray | cirq.STATE_VECTOR_LIKE = 0, dtype: type[np.complexfloating] = np.complex64, classical_data: cirq.ClassicalDataStore | None = None, # should_preserve_initial_state: bool = True, ):Then in _BufferedStateVector.create():
@classmethod def create( cls, *, initial_state: np.ndarray | cirq.STATE_VECTOR_LIKE = 0, qid_shape: tuple[int, ...] | None = None, dtype: type[np.complexfloating] | None = None, buffer: np.ndarray | None = None, # should_preserve_initial_state: bool = True, ): # some code here # .... if np.may_share_memory(state_vector, initial_state) and should_preserve_initial_state: # listen to the flag state_vector = state_vector.copy() # ....These code changes would appear in Cirq, not qsim, but I believe may resolve the issue that the internal user was experiencing.
Thank you for the proposed solution. In addition to copying the initial state,
_BufferedStateVectorcreates a buffer in__init__. The buffer creation should also be toggled by this flag.Reacted by Jagan Kalsi@sergeisakov - are simulations of 40-qubit state vectors even feasible in qsim?
On my cloudtop with 94GB of memory I am not even able to allocate such state vector once as
np.zeros(2**40, dtype=np.int8)fails with the same_ArrayMemoryErroras in the description. I can dox = np.arange(2**30, dtype=float), but thenx.dot(x)takes about 160ms. For a 40-qubit state vector the dot product can be projected to take 1000 times more, i.e., almost 3 minutes. That sounds painstakingly slow as I'd expect there would be more full-state vector operations in a simulation run.@mhucka - can you please check with the user for an exact example of that simulation or - if confidential - for a similar sized textbook case? I'd be curious if it is even feasible to handle 40 qubits with a C++ without an extra overhead for Python / numpy interfaces.
Yes, simulations of 40-qubit state vectors are feasible with qsim. I ran such simulations. It requires 8 TB of memory. For instance, there are x4 machines in Google Cloud that have up to 32 TB. That's enough for 41 qubits.
Yes, simulations of 40-qubit state vectors are feasible with qsim. I ran such simulations. It requires 8 TB of memory. For instance, there are x4 machines in Google Cloud that have up to 32 TB. That's enough for 41 qubits.
@mhucka - can you please confirm with the reporting user if the failing example was run on a similar machine with sufficient memory?
Reacted by Michael Hucka@mhucka - can you please confirm with the reporting user if the failing example was run on a similar machine with sufficient memory?
I've contacted them and will write a note here when I hear back.
TODO - test the example with Cirq PR quantumlib/Cirq#8054 making sure the new flag for memory reuse is on.
Update from Cirq Cynq 2026-05-13: we verified with the user that they had at least 8 TB available on the system (but probably not 16 TB), and have requested an example that can be used to reproduce the failure and be used to test before/after behavior of this PR.
We probably don't need 40 qubits to reproduce this issue and test the before/after behavior. For example, 33 qubits only take 64 GB, so it should run fine on a 96 GB machine.
Reacted by Michael HuckaTODO - test the example with Cirq PR quantumlib/Cirq#8054 making sure the new flag for memory reuse is on.
@pavoljuhas I was thinking how to demonstrate a difference being made with the addition of the
should_preserve_initial_stateflag...I ran a short script confirming the efficacy of the flag on my local machine (MacOS 16gb RAM).
The script spawned subprocesses as following, testing qubit counts n= [25-30], and flag=true/false:
import sys import numpy as np from cirq.sim.state_vector_simulation_state import _BufferedStateVector n = {n} flag = {flag} # Fully-written state vector: every amplitude physically backed # in RAM because vals are assigned via rng.random(...), # matching what qsim may produce after a real simulation rng = np.random.default_rng(5) flat = rng.random(2 * 2**n, dtype=np.float32).view(np.complex64) flat /= np.linalg.norm(flat) sv = flat.reshape(tuple([2] * n)) try: bsv = _BufferedStateVector.create( initial_state=sv, qid_shape=tuple([2] * n), dtype=np.complex64, should_preserve_initial_state=flag, ) print("SUCCEEDED") sys.exit(0) except MemoryError as exc: print(f"OOM: {{exc}}") sys.exit(42)Then measured their execution time, with a timeout of 60s to catch automatic disk-swapping when out of RAM (which MacOS does). I tested the creation of BufferedStateVector instances and got the following results:
Total RAM : 16.0 GB Available : 7.8 GB Timeout : 60s per test Range (# of qubits) : n = 25 to 30 n sv size flag=True flag=False --------------------------------------------------------------- 25 0.2 GB ok (1.62s) ok (1.55s) 26 0.5 GB ok (1.73s) ok (1.66s) 27 1.0 GB ok (2.30s) ok (2.26s) 28 2.0 GB ok (3.47s) ok (3.21s) 29 4.0 GB ok (8.93s) ok (5.12s) 30 8.0 GB ok (41.64s) ok (9.33s)At 30 qubits, when the flag was set to True, the disk-swapping time overhead is evident. The copy of the state vector occurred, and the process consumed all available RAM and had to resort to using disk memory. However, when the flag was set to False, the execution time greatly decreased, as expected, indicating that by avoiding the statevector copy, the process did not run out of RAM.
This suggests the flag is effective - at least as an example on a smaller system, as suggested by @sergeisakov. On a machine without automatic disk-swap support (like the original user, probably?), I suspect the failure would have looked like the original reported stack trace, instead of a massive increase in execution time.
If I am missing or misunderstanding something, please let me know. I can provide the full python script if anyone would like to run and test in their own environment, too.
Update from Cirq Cynq 2026-05-13: we verified with the user that they had at least 8 TB available on the system (but probably not 16 TB), and have requested an example that can be used to reproduce the failure and be used to test before/after behavior of this PR.
To close this loop: the user does not recall the specific simulation they were doing, but says that it should be demonstrable with any suitably large circuit.
I think it really sounds the problem was hitting the memory limit on the system they were using. (Probably an 8 TB system, per earlier comments.)
Reacted by Jagan Kalsi@jkalsi1 - AFAICT, the check in #893 (comment) merely verifies that a delayed allocation of the
_BufferedStateVector._bufferis indeed delayed. It does not tell if there is any difference for memory use during theQSimSimulator.simulatecall as in the issue description above.- added a commit that references this issue
on Sep 4, 2026
What is your request or suggestion?
An internal user reported getting the following error with a 40 qubit simulation:
They were able to work around the problem by calling
QSimSimulator()._simulate_impl()directly.What solution or approach do you envision?
We should investigate whether the copy can be handled in a more user-friendly way.
How urgent is this for you?
P2 – needed within two quarters