Skip to content
VVeiCaoPublic

About

4Director: Controlling Video World Models with Rigid 3D Geometry

Resources

Stars

22 stars

Watchers

0 watching

Forks

Repository files navigation

4Director

4Director: Controlling Video World Models with Rigid 3D Geometry

arXiv Project Page

⭐ If you find 4Director useful, please give this repo a star. Thank you!

Wei Cao1,2   Hao Zhang2   Vikram Voleti1   Yuqun Wu2,1   Mallikarjun B R1   Shimon Vainer1   Mark Boss1   Yaoyao Liu2

1Stability AI     2University of Illinois at Urbana-Champaign

Stability AI          University of Illinois at Urbana-Champaign

Object control      New object insertion      Camera control

TL;DR: A video world model conditioned on an explicit 4D scene representation, giving direct 3D control over camera and multi-object motion. One input image becomes an editable 4D scene, and every control below is a rigid 3D transform applied inside that scene.

Input image
Input image
Reference photo for the inserted camel
New object reference
Controls Scene point cloud — click to orbit Generated video
Object control View the interactive Hiker point cloud Hiker generated video
Object control New object insertion View the interactive Hiker and Camel point cloud Hiker and Camel generated video
Object control New object insertion
Camera control
View the interactive Hiker, Camel and camera point cloud Hiker, Camel and camera generated video

Installation

Validated on Linux with NVIDIA H200 GPUs. The whole pipeline runs on one GPU with at least 80 GB of memory, and the weights and model caches take about 110 GB of disk.

The environment pins Python 3.11 + PyTorch 2.7.1 + xformers 0.0.31.post1 + CUDA 12.8 (cu128 wheels) + cuDNN 9.13.1, so the NVIDIA driver must support CUDA 12.8. Compiling the extensions also needs the CUDA 12.8 toolkit (nvcc 12.8; a 13.x toolkit does not work). install_extensions.sh uses $CUDA_HOME if it is set, else /usr/local/cuda-12.8, else the nvcc on PATH.

Step 1 — Create the environment

# Clone the repository and its submodules.
git clone --recurse-submodules https://github.com/VVeiCao/4Director.git
cd 4Director
# Cloned without --recurse-submodules? Fetch them now:
# git submodule update --init --recursive

# Create the conda environment and install 4Director.
conda env create -f environments/4director.yml
conda activate 4director
pip install -e '.[ui]'

# Build PyTorch3D, Pixal3D's CUDA extensions, and MegaSAM, and add cuDNN 9.13.1 (needs a GPU and the CUDA 12.8 toolkit).
bash environments/install_extensions.sh

Step 2 — Download the weights

# SAM 2, Wan2.1-VACE-14B, and the 4Director Motion Adapter, into ./models.
python download_models.py

The app's first Generate 3D also fetches MoGe-2, UniDepth, Qwen3-VL, and Pixal3D, about 31 GB, into ./models/cache. If some weights already live elsewhere, copy components.env.example to components.env and point it at them; shell exports of those names are ignored.

Reproduce the Teaser Cases

# Generate the three teaser videos into output/<case>/output.mp4.
python run.py --config configs/cases/object-motion.yaml     # Hiker
python run.py --config configs/cases/object-insertion.yaml  # Hiker + Camel
python run.py --config configs/cases/joint-camera.yaml      # Hiker + Camel + camera
# Play a Case in 3D at http://localhost:8080: the point cloud, with the objects and camera moving along their paths.
python run.py viser --config configs/cases/joint-camera.yaml

The joint-camera Case in the Viser viewer

Generate from a Single Image

Step 1 — Build the 3D scene

python apps/frame0_gradio.py --port 7860

Open http://localhost:7860; the terminal lists anything the app cannot find yet. On a remote GPU machine, forward the port first (ssh -L 7860:localhost:7860 <host>), and port 8080 the same way for Viser.

Upload, click, mask, and 3D reconstruction in the app

  1. Upload an 832×480 image, or keep the golden retriever example. Other sizes are refused rather than stretched.
  2. Click the subject to move. Negative clicks remove anything SAM2 picks up by mistake.
  3. Generate the mask and confirm it. Nothing expensive runs before this.
  4. Generate 3D builds the object mesh, the background point cloud, and an editable caption.

Step 2 — Choose the motion

Motion presets with the 3D scene and preview clips

Pick one of six motions (turn or move the object, or orbit the camera) and preview it before generating.

Step 3 — Render and generate the video

Run the two commands printed at the bottom of the page:

# Render the depth control for the chosen motion.
python run.py stage3 --config output/frame0-demo/<run-id>/case.yaml \
    --run-root output/frame0-demo/<run-id> --force

# Generate the video into output/frame0-demo/<run-id>/output-seed123.mp4.
python run.py stage4 --config output/frame0-demo/<run-id>/case.yaml \
    --run-root output/frame0-demo/<run-id> --checkpoint models/4Director/step-2610.safetensors \
    --seed 123 --force

If a video looks off, rerun only the second command with another --seed; each seed writes its own output-seed<seed>.mp4, and the control does not depend on it. Runs are saved under output/frame0-demo/ and can be reopened, mask, mesh, and caption included, from the list at the top of the page.

License and Attribution

The code and model weights are released under the Stability AI Community License: free for research, non-commercial, and commercial use by organizations and individuals with annual revenue up to US $1,000,000. Above that, commercial use needs an Enterprise License from Stability AI. See the Hugging Face model card for details.

The Git submodules under third_party/, the extensions install_extensions.sh builds, and the models downloaded at run time keep their own licenses. Three of them allow research and other non-commercial use only: UniDepth (CC BY-NC 4.0), which builds the scene point cloud, and nvdiffrast and nvdiffrec (NVIDIA Source Code License), which Pixal3D renders with. Reproducing the teaser Cases runs none of them; building a new scene or object mesh does. Pixal3D's DINOv3 encoder comes under the DINOv3 License.

Citation

If you find 4Director useful, please cite:

@misc{cao20264directorcontrollingvideoworld,
  title={4Director: Controlling Video World Models with Rigid 3D Geometry},
  author={Wei Cao and Hao Zhang and Vikram Voleti and Yuqun Wu and Mallikarjun B R and Shimon Vainer and Mark Boss and Yaoyao Liu},
  year={2026},
  eprint={2610.02160},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2610.02160},
}

Acknowledgements

We thank the authors of Wan2.1, VACE, DiffSynth-Studio, MegaSaM, MoGe-2, UniDepthV2, Pixal3D, TRELLIS.2, TAPIP3D, SAM 2, SAM 3, Qwen3-VL, PyTorch3D, Viser, RealCOD-25K, and DAVIS for releasing their code, models, and data.

About

4Director: Controlling Video World Models with Rigid 3D Geometry

Resources

Stars

22 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages