Long-Horizon Collaboration of Multi-agent LLMs
Benchmark your model. Build your agent harness. Watch a team work together.
Website & leaderboard · Paper · Quickstart · Custom agents · Submit results
AgentWorld is a 2D multiplayer environment for evaluating AI agents. Agents interact with a shared game world through tools: gathering resources, crafting, fighting, communicating, and coordinating toward task objectives.
This repository brings together the game engine, Python reference harness, versioned tasks, and evaluation utilities so you can follow an experiment from task → actions → trajectory → score.
Ten agents, distinct roles, one shared world. Figure 1 from Shu et al., AgentWorld (CC BY 4.0).
| I want to… | Start here |
|---|---|
| Benchmark a model | Use an OpenAI-compatible endpoint with the reference runner. |
| Benchmark my agent harness | Bring your planner, memory, or framework through a custom adapter. |
| Understand a run | Inspect actions and observations with trajectory visualizations. |
| Contribute or review results | Read the contributor guide and submission requirements. |
The environment, agent interaction model, and representative tools. Figure 2 from Shu et al., AgentWorld (CC BY 4.0). Click either figure for full resolution.
You need Python 3.10+, a running AgentWorld game instance, and a model endpoint supporting the reference adapter's tool-calling contract. API-based models do not need a local GPU. All commands below run from the repository root.
git clone --branch develop https://github.com/openagents-org/agentworld.git
cd agentworld
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -r agents/requirements.txtStart a dedicated game instance using the server setup guide. The game HTTP API and your model endpoint are two separate services.
export AGENTWORLD_BASE_URL='http://localhost:7031'
export MODEL_BASE_URL='http://localhost:8000/v1'
export MODEL_NAME='your-tool-calling-model-id'
export MODEL_API_KEY='your-model-api-key'Use MODEL_API_KEY=unused for an unauthenticated local model endpoint. The
example configuration reads
these variables from your shell. See the quickstart
for endpoint compatibility and authentication options.
python agents/run.py \
--task data_v0.1_multi/v1.3_benchmark/task_01_magic_staff.yaml \
--agent agents/configs/openai-compatible.example.yaml \
--output runs/my-model/smoke \
--no-split-screenOpen runs/my-model/smoke/task_01_trajectory.json to inspect the recorded rounds,
actions, observations, and verification metrics. Once the smoke task runs
correctly, move to a complete suite.
Run the main suite and calculate success rate
Use a fresh directory for each model, harness configuration, and trial.
python agents/run.py \
--task-folder data_v0.1_multi/v1.3_benchmark \
--agent agents/configs/openai-compatible.example.yaml \
--output runs/my-model/main/trial-1 \
--no-split-screen
python benchmarks/score.py \
--suite main \
--trajectories runs/my-model/main/trial-1 \
--output runs/my-model/main/trial-1/scores.jsonFor augmented runs, use data_v0.1_multi/v1.3_augmented, a separate output
directory, and --suite augmented when scoring. Do not use the parent
data_v0.1_multi/ directory as a suite: it includes historical datasets.
The runner can overwrite results in reused directories. Give concurrent experiments separate game instances to avoid shared-state interference.
Model integration: point the reference configuration at your model endpoint. Harness integration: start from the custom Python adapter and example configuration.
# Uses the same model and game environment variables as above.
PYTHONPATH=examples/agents python agents/run.py \
--task data_v0.1_multi/v1.3_benchmark/task_01_magic_staff.yaml \
--agent agents/configs/custom.example.yaml \
--output runs/my-harness/smoke \
--no-split-screenThe example initially preserves the reference policy so you can verify the integration before changing behavior. The custom-agent guide explains the turn contract, trajectory records, and how to connect an independent harness through the game API.
| Metric | What is available here |
|---|---|
| Success rate · SR | Offline, task-specific verification with coverage, errors, and per-task outcomes. Full-suite SR is reported only when coverage is complete and there are no scoring errors. |
| Causal Collaboration Effectiveness · CCE | Research analysis using an LLM judge. Record the judge, configuration, and protocol alongside results. |
| Partial success rate · PSR | A validated implementation reproducing the paper's metric is not yet available in this checkout. |
The checkout contains 200 augmented task files; the paper reports experiments on 100 augmented variants. Record the exact task manifest when comparing results.
See scoring and protocol notes before comparing results. A partial run's observed SR is not a full-suite score; custom harnesses, modified budgets, and alternative judges must be identified in submissions.
Ready to share? Follow the submission guide and open a result-submission GitHub issue. Include your model and harness versions, task revision, settings, trials, scores, and raw trajectories. Reviewers check the evidence before leaderboard publication.
agentworld/
├── agents/ Reference runner, model adapters, tools, and configs
├── benchmarks/ Offline scoring and benchmark entry-point docs
├── data_v0.1_multi/ Versioned multi-agent tasks and task verifiers
├── data_v0.1_solo/ Solo task definitions
├── examples/agents/ Custom-harness integration example
├── packages/ TypeScript game server, browser client, and shared code
├── analysis/ Research analysis and historical report archive
├── scripts/ Development utilities and manual diagnostics
├── docs/ Benchmark guides and game documentation
└── tests/benchmark/ Offline onboarding regression tests
The repository map distinguishes the supported onboarding
path from historical experiments. Generated runs and visualizations belong in
ignored runs/ directories. The leaderboard website lives in the separate
agentworld-web repository.
- Learn the world: game tools, crafting, and regions.
- Customize experiments: prompt templates and baseline flags.
- Improve the benchmark: contribute adapters, task/verifier fixes, documentation, or submission reviews. Start with CONTRIBUTING.md.
- Revisit earlier work: see legacy usage and the research archive.
Built upon Kaetram, which expands on Little Workshop's BrowserQuest. Code is licensed under MPL-2.0. Paper figures are by Shu et al., licensed under CC BY 4.0; see figure sources and attribution.