⭐ If you find 4Director useful, please give this repo a star. Thank you!
Wei Cao1,2 Hao Zhang2 Vikram Voleti1 Yuqun Wu2,1 Mallikarjun B R1 Shimon Vainer1 Mark Boss1 Yaoyao Liu2
1Stability AI 2University of Illinois at Urbana-Champaign
Object control
New object insertion
Camera control
TL;DR: A video world model conditioned on an explicit 4D scene representation, giving direct 3D control over camera and multi-object motion. One input image becomes an editable 4D scene, and every control below is a rigid 3D transform applied inside that scene.
| Controls | Scene point cloud — |
Generated video |
|---|---|---|
|
|
![]() |
![]() |
|
|
![]() |
![]() |
|
|
![]() |
![]() |
Validated on Linux with NVIDIA H200 GPUs. The whole pipeline runs on one GPU with at least 80 GB of memory, and the weights and model caches take about 110 GB of disk.
The environment pins Python 3.11 + PyTorch 2.7.1 + xformers 0.0.31.post1 +
CUDA 12.8 (cu128 wheels) + cuDNN 9.13.1, so the NVIDIA driver must support
CUDA 12.8.
Compiling the extensions also needs the CUDA 12.8 toolkit (nvcc 12.8; a
13.x toolkit does not work). install_extensions.sh uses $CUDA_HOME if it is
set, else /usr/local/cuda-12.8, else the nvcc on PATH.
# Clone the repository and its submodules.
git clone --recurse-submodules https://github.com/VVeiCao/4Director.git
cd 4Director
# Cloned without --recurse-submodules? Fetch them now:
# git submodule update --init --recursive
# Create the conda environment and install 4Director.
conda env create -f environments/4director.yml
conda activate 4director
pip install -e '.[ui]'
# Build PyTorch3D, Pixal3D's CUDA extensions, and MegaSAM, and add cuDNN 9.13.1 (needs a GPU and the CUDA 12.8 toolkit).
bash environments/install_extensions.sh# SAM 2, Wan2.1-VACE-14B, and the 4Director Motion Adapter, into ./models.
python download_models.pyThe app's first Generate 3D also fetches MoGe-2, UniDepth, Qwen3-VL, and
Pixal3D, about 31 GB, into ./models/cache. If some weights already live
elsewhere, copy components.env.example to components.env and point it at
them; shell exports of those names are ignored.
# Generate the three teaser videos into output/<case>/output.mp4.
python run.py --config configs/cases/object-motion.yaml # Hiker
python run.py --config configs/cases/object-insertion.yaml # Hiker + Camel
python run.py --config configs/cases/joint-camera.yaml # Hiker + Camel + camera# Play a Case in 3D at http://localhost:8080: the point cloud, with the objects and camera moving along their paths.
python run.py viser --config configs/cases/joint-camera.yamlpython apps/frame0_gradio.py --port 7860Open http://localhost:7860; the terminal lists anything the app cannot find
yet. On a remote GPU machine, forward the port first
(ssh -L 7860:localhost:7860 <host>), and port 8080 the same way for Viser.
- Upload an 832×480 image, or keep the golden retriever example. Other sizes are refused rather than stretched.
- Click the subject to move. Negative clicks remove anything SAM2 picks up by mistake.
- Generate the mask and confirm it. Nothing expensive runs before this.
- Generate 3D builds the object mesh, the background point cloud, and an editable caption.
Pick one of six motions (turn or move the object, or orbit the camera) and preview it before generating.
Run the two commands printed at the bottom of the page:
# Render the depth control for the chosen motion.
python run.py stage3 --config output/frame0-demo/<run-id>/case.yaml \
--run-root output/frame0-demo/<run-id> --force
# Generate the video into output/frame0-demo/<run-id>/output-seed123.mp4.
python run.py stage4 --config output/frame0-demo/<run-id>/case.yaml \
--run-root output/frame0-demo/<run-id> --checkpoint models/4Director/step-2610.safetensors \
--seed 123 --forceIf a video looks off, rerun only the second command with another --seed;
each seed writes its own output-seed<seed>.mp4, and the control does not
depend on it. Runs are saved under output/frame0-demo/ and can be reopened,
mask, mesh, and caption included, from the list at the top of the page.
The code and model weights are released under the Stability AI Community License: free for research, non-commercial, and commercial use by organizations and individuals with annual revenue up to US $1,000,000. Above that, commercial use needs an Enterprise License from Stability AI. See the Hugging Face model card for details.
The Git submodules under third_party/, the extensions install_extensions.sh
builds, and the models downloaded at run time keep their own licenses. Three of
them allow research and other non-commercial use only: UniDepth (CC BY-NC 4.0),
which builds the scene point cloud, and nvdiffrast and nvdiffrec (NVIDIA Source
Code License), which Pixal3D renders with. Reproducing the teaser Cases runs none
of them; building a new scene or object mesh does. Pixal3D's DINOv3 encoder comes
under the DINOv3 License.
If you find 4Director useful, please cite:
@misc{cao20264directorcontrollingvideoworld,
title={4Director: Controlling Video World Models with Rigid 3D Geometry},
author={Wei Cao and Hao Zhang and Vikram Voleti and Yuqun Wu and Mallikarjun B R and Shimon Vainer and Mark Boss and Yaoyao Liu},
year={2026},
eprint={2610.02160},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2610.02160},
}We thank the authors of Wan2.1, VACE, DiffSynth-Studio, MegaSaM, MoGe-2, UniDepthV2, Pixal3D, TRELLIS.2, TAPIP3D, SAM 2, SAM 3, Qwen3-VL, PyTorch3D, Viser, RealCOD-25K, and DAVIS for releasing their code, models, and data.











