Skip to content

Provide a way to avoid copying state vector from Cirq #893

Description

@mhucka

What is your request or suggestion?

An internal user reported getting the following error with a 40 qubit simulation:

Traceback (most recent call last):
  File "/tmp/test/exact_sim.py", line 23, in <module>
    psi = qsim.simulate(circuit, qubit_order=qubits).state_vector()
          ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/tmp/venv/lib64/python3.12/site-packages/cirq/sim/simulator.py", line 496, in simulate
    return self.simulate_sweep(
           ^^^^^^^^^^^^^^^^^^^^
  File "/tmp/venv/lib64/python3.12/site-packages/cirq/sim/simulator.py", line 511, in simulate_sweep
    return list(self.simulate_sweep_iter(program, params, qubit_order, initial_state))
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/tmp/venv/lib64/python3.12/site-packages/qsimcirq/qsim_simulator.py", line 560, in simulate_sweep_iter
    final_state = cirq.StateVectorSimulationState(
                  ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/tmp/venv/lib64/python3.12/site-packages/cirq/sim/state_vector_simulation_state.py", line 351, in __init__
    state = _BufferedStateVector.create(
            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/tmp/venv/lib64/python3.12/site-packages/cirq/sim/state_vector_simulation_state.py", line 84, in create
    state_vector = state_vector.copy()
                   ^^^^^^^^^^^^^^^^^^^
numpy._core._exceptions._ArrayMemoryError: Unable to allocate 8.00 TiB for an array with shape (2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2) and data type complex64

They were able to work around the problem by calling QSimSimulator()._simulate_impl() directly.

What solution or approach do you envision?

We should investigate whether the copy can be handled in a more user-friendly way.

How urgent is this for you?

P2 – needed within two quarters

Activity

  1. added
    a: qsimcirqInvolves the qsim-Cirq integration layer
    a: pythonInvolves the Python code in qsim
    a: performanceInvolves performance problems or improvements
    P2Priority: medium priority
    on Sep 23, 2025
  2. modified the milestone: Release 0.24.0 on Sep 23, 2025
  3. jkalsi1 commented on Apr 2, 2026

    @jkalsi1
    Member

    Hi @mhucka . In cirq-core/cirq/sim/state_vector_simulation_state.py.__BufferedStateVector.create(), the following condition evaluates truthy, which forces the copy of the array to be made, causing the above error:

    if np.may_share_memory(state_vector, initial_state):
            state_vector = state_vector.copy()
    

    It appears to me that this is to protect the integrity of the original initial_state argument passed to _BufferedStateVector.create() against unwanted mutations.

    Could we create a flag in cirq-core/cirq/sim/state_vector_simulation_state.py.StateVectorSimulationState that allows the caller to choose whether they care about the integrity of the initial_state array or not. If they don't care, then skip the copy condition.

    Something like:

    In StateVectorSimulationState:

        def __init__(
            self,
            *,
            available_buffer: np.ndarray | None = None,
            prng: np.random.RandomState | None = None,
            qubits: Sequence[cirq.Qid] | None = None,
            initial_state: np.ndarray | cirq.STATE_VECTOR_LIKE = 0,
            dtype: type[np.complexfloating] = np.complex64,
            classical_data: cirq.ClassicalDataStore | None = None,
            # should_preserve_initial_state: bool = True,
        ):
    

    Then in _BufferedStateVector.create():

        @classmethod
        def create(
            cls,
            *,
            initial_state: np.ndarray | cirq.STATE_VECTOR_LIKE = 0,
            qid_shape: tuple[int, ...] | None = None,
            dtype: type[np.complexfloating] | None = None,
            buffer: np.ndarray | None = None,
            # should_preserve_initial_state: bool = True,
        ):
    
      # some code here 
     # ....
    if np.may_share_memory(state_vector, initial_state) and should_preserve_initial_state: # listen to the flag
            state_vector = state_vector.copy()
    # ....
    

    These code changes would appear in Cirq, not qsim, but I believe may resolve the issue that the internal user was experiencing.

  4. jkalsi1 commented on Apr 15, 2026

    @jkalsi1
    Member
  5. sergeisakov commented on Apr 20, 2026

    @sergeisakov
    Collaborator

    Thank you for the proposed solution. In addition to copying the initial state, _BufferedStateVector creates a buffer in __init__. The buffer creation should also be toggled by this flag.

  6. pavoljuhas commented on May 12, 2026

    @pavoljuhas
    Collaborator

    @sergeisakov - are simulations of 40-qubit state vectors even feasible in qsim?

    On my cloudtop with 94GB of memory I am not even able to allocate such state vector once as np.zeros(2**40, dtype=np.int8) fails with the same _ArrayMemoryError as in the description. I can do x = np.arange(2**30, dtype=float), but then x.dot(x) takes about 160ms. For a 40-qubit state vector the dot product can be projected to take 1000 times more, i.e., almost 3 minutes. That sounds painstakingly slow as I'd expect there would be more full-state vector operations in a simulation run.

    @mhucka - can you please check with the user for an exact example of that simulation or - if confidential - for a similar sized textbook case? I'd be curious if it is even feasible to handle 40 qubits with a C++ without an extra overhead for Python / numpy interfaces.

  7. sergeisakov commented on May 12, 2026

    @sergeisakov
    Collaborator

    Yes, simulations of 40-qubit state vectors are feasible with qsim. I ran such simulations. It requires 8 TB of memory. For instance, there are x4 machines in Google Cloud that have up to 32 TB. That's enough for 41 qubits.

  8. pavoljuhas commented on May 12, 2026

    @pavoljuhas
    Collaborator

    Yes, simulations of 40-qubit state vectors are feasible with qsim. I ran such simulations. It requires 8 TB of memory. For instance, there are x4 machines in Google Cloud that have up to 32 TB. That's enough for 41 qubits.

    @mhucka - can you please confirm with the reporting user if the failing example was run on a similar machine with sufficient memory?

  9. mhucka commented on May 13, 2026

    @mhucka
    CollaboratorAuthor

    @mhucka - can you please confirm with the reporting user if the failing example was run on a similar machine with sufficient memory?

    I've contacted them and will write a note here when I hear back.

  10. pavoljuhas commented on May 13, 2026

    @pavoljuhas
    Collaborator

    TODO - test the example with Cirq PR quantumlib/Cirq#8054 making sure the new flag for memory reuse is on.

  11. mhucka commented on May 13, 2026

    @mhucka
    CollaboratorAuthor

    Update from Cirq Cynq 2026-05-13: we verified with the user that they had at least 8 TB available on the system (but probably not 16 TB), and have requested an example that can be used to reproduce the failure and be used to test before/after behavior of this PR.

  12. sergeisakov commented on May 13, 2026

    @sergeisakov
    Collaborator

    We probably don't need 40 qubits to reproduce this issue and test the before/after behavior. For example, 33 qubits only take 64 GB, so it should run fine on a 96 GB machine.

  13. jkalsi1 commented on May 15, 2026

    @jkalsi1
    Member

    TODO - test the example with Cirq PR quantumlib/Cirq#8054 making sure the new flag for memory reuse is on.

    @pavoljuhas I was thinking how to demonstrate a difference being made with the addition of the should_preserve_initial_state flag...

    I ran a short script confirming the efficacy of the flag on my local machine (MacOS 16gb RAM).

    The script spawned subprocesses as following, testing qubit counts n= [25-30], and flag=true/false:

        import sys
        import numpy as np
        from cirq.sim.state_vector_simulation_state import _BufferedStateVector
    
        n = {n}
        flag = {flag}
    
        # Fully-written state vector: every amplitude physically backed
        # in RAM because vals are assigned via rng.random(...),
        # matching what qsim may produce after a real simulation
        rng = np.random.default_rng(5)
        flat = rng.random(2 * 2**n, dtype=np.float32).view(np.complex64)
        flat /= np.linalg.norm(flat)
        sv = flat.reshape(tuple([2] * n))
    
        try:
            bsv = _BufferedStateVector.create(
                initial_state=sv,
                qid_shape=tuple([2] * n),
                dtype=np.complex64,
                should_preserve_initial_state=flag,
            )
            print("SUCCEEDED")
            sys.exit(0)
        except MemoryError as exc:
            print(f"OOM: {{exc}}")
            sys.exit(42)
    

    Then measured their execution time, with a timeout of 60s to catch automatic disk-swapping when out of RAM (which MacOS does). I tested the creation of BufferedStateVector instances and got the following results:

    Total RAM           : 16.0 GB
    Available           : 7.8 GB
    Timeout             : 60s per test
    Range (# of qubits)    : n = 25 to 30
    
      n   sv size               flag=True              flag=False
    ---------------------------------------------------------------
     25     0.2 GB    ok           (1.62s)    ok           (1.55s)
     26     0.5 GB    ok           (1.73s)    ok           (1.66s)
     27     1.0 GB    ok           (2.30s)    ok           (2.26s)
     28     2.0 GB    ok           (3.47s)    ok           (3.21s)
     29     4.0 GB    ok           (8.93s)    ok           (5.12s)
     30     8.0 GB   ok           (41.64s)    ok           (9.33s)
    

    At 30 qubits, when the flag was set to True, the disk-swapping time overhead is evident. The copy of the state vector occurred, and the process consumed all available RAM and had to resort to using disk memory. However, when the flag was set to False, the execution time greatly decreased, as expected, indicating that by avoiding the statevector copy, the process did not run out of RAM.



    This suggests the flag is effective - at least as an example on a smaller system, as suggested by @sergeisakov. On a machine without automatic disk-swap support (like the original user, probably?), I suspect the failure would have looked like the original reported stack trace, instead of a massive increase in execution time.

    If I am missing or misunderstanding something, please let me know. I can provide the full python script if anyone would like to run and test in their own environment, too.

  14. mhucka commented on May 15, 2026

    @mhucka
    CollaboratorAuthor

    Update from Cirq Cynq 2026-05-13: we verified with the user that they had at least 8 TB available on the system (but probably not 16 TB), and have requested an example that can be used to reproduce the failure and be used to test before/after behavior of this PR.

    To close this loop: the user does not recall the specific simulation they were doing, but says that it should be demonstrable with any suitably large circuit.

    I think it really sounds the problem was hitting the memory limit on the system they were using. (Probably an 8 TB system, per earlier comments.)

  15. pavoljuhas commented on May 28, 2026

    @pavoljuhas
    Collaborator

    @jkalsi1 - AFAICT, the check in #893 (comment) merely verifies that a delayed allocation of the _BufferedStateVector._buffer is indeed delayed. It does not tell if there is any difference for memory use during the QSimSimulator.simulate call as in the issue description above.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    P2Priority: medium prioritya: performanceInvolves performance problems or improvementsa: pythonInvolves the Python code in qsima: qsimcirqInvolves the qsim-Cirq integration layercontributors welcomeHelp with this would be appreciatedno QC knowledge neededDoes not require knowledge of quantum computing

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions