Skip to content

About

No description, website, or topics provided.

Resources

Contributing

Stars

27 stars

Watchers

1 watching

Forks

Latest commit

 

History

3,157 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

AgentWorld

Long-Horizon Collaboration of Multi-agent LLMs

Benchmark your model. Build your agent harness. Watch a team work together.

Website & leaderboard · Paper · Quickstart · Custom agents · Submit results

License: MPL-2.0 Runner: Python 3.10 or newer Main suite: 100 tasks Augmented suite: 200 variants

AgentWorld is a 2D multiplayer environment for evaluating AI agents. Agents interact with a shared game world through tools: gathering resources, crafting, fighting, communicating, and coordinating toward task objectives.

This repository brings together the game engine, Python reference harness, versioned tasks, and evaluation utilities so you can follow an experiment from task → actions → trajectory → score.

Ten AgentWorld characters with crafting, mining, combat, and support roles coordinating through live chat.
Ten agents, distinct roles, one shared world. Figure 1 from Shu et al., AgentWorld (CC BY 4.0).

Choose your starting point

I want to… Start here
Benchmark a model Use an OpenAI-compatible endpoint with the reference runner.
Benchmark my agent harness Bring your planner, memory, or framework through a custom adapter.
Understand a run Inspect actions and observations with trajectory visualizations.
Contribute or review results Read the contributor guide and submission requirements.

Inside the benchmark

AgentWorld overview: a sandbox with diverse biomes, agents communicating without seeing one another’s internal state, and high-level game tools.
The environment, agent interaction model, and representative tools. Figure 2 from Shu et al., AgentWorld (CC BY 4.0). Click either figure for full resolution.

Your first experiment

You need Python 3.10+, a running AgentWorld game instance, and a model endpoint supporting the reference adapter's tool-calling contract. API-based models do not need a local GPU. All commands below run from the repository root.

1 · Install the runner

git clone --branch develop https://github.com/openagents-org/agentworld.git
cd agentworld
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -r agents/requirements.txt

2 · Connect the game and model

Start a dedicated game instance using the server setup guide. The game HTTP API and your model endpoint are two separate services.

export AGENTWORLD_BASE_URL='http://localhost:7031'
export MODEL_BASE_URL='http://localhost:8000/v1'
export MODEL_NAME='your-tool-calling-model-id'
export MODEL_API_KEY='your-model-api-key'

Use MODEL_API_KEY=unused for an unauthenticated local model endpoint. The example configuration reads these variables from your shell. See the quickstart for endpoint compatibility and authentication options.

3 · Run one task

python agents/run.py \
  --task data_v0.1_multi/v1.3_benchmark/task_01_magic_staff.yaml \
  --agent agents/configs/openai-compatible.example.yaml \
  --output runs/my-model/smoke \
  --no-split-screen

Open runs/my-model/smoke/task_01_trajectory.json to inspect the recorded rounds, actions, observations, and verification metrics. Once the smoke task runs correctly, move to a complete suite.

Run the main suite and calculate success rate

Use a fresh directory for each model, harness configuration, and trial.

python agents/run.py \
  --task-folder data_v0.1_multi/v1.3_benchmark \
  --agent agents/configs/openai-compatible.example.yaml \
  --output runs/my-model/main/trial-1 \
  --no-split-screen

python benchmarks/score.py \
  --suite main \
  --trajectories runs/my-model/main/trial-1 \
  --output runs/my-model/main/trial-1/scores.json

For augmented runs, use data_v0.1_multi/v1.3_augmented, a separate output directory, and --suite augmented when scoring. Do not use the parent data_v0.1_multi/ directory as a suite: it includes historical datasets.

The runner can overwrite results in reused directories. Give concurrent experiments separate game instances to avoid shared-state interference.

Bring your own harness

Model integration: point the reference configuration at your model endpoint. Harness integration: start from the custom Python adapter and example configuration.

# Uses the same model and game environment variables as above.
PYTHONPATH=examples/agents python agents/run.py \
  --task data_v0.1_multi/v1.3_benchmark/task_01_magic_staff.yaml \
  --agent agents/configs/custom.example.yaml \
  --output runs/my-harness/smoke \
  --no-split-screen

The example initially preserves the reference policy so you can verify the integration before changing behavior. The custom-agent guide explains the turn contract, trajectory records, and how to connect an independent harness through the game API.

Scores you can inspect

Metric What is available here
Success rate · SR Offline, task-specific verification with coverage, errors, and per-task outcomes. Full-suite SR is reported only when coverage is complete and there are no scoring errors.
Causal Collaboration Effectiveness · CCE Research analysis using an LLM judge. Record the judge, configuration, and protocol alongside results.
Partial success rate · PSR A validated implementation reproducing the paper's metric is not yet available in this checkout.

The checkout contains 200 augmented task files; the paper reports experiments on 100 augmented variants. Record the exact task manifest when comparing results.

See scoring and protocol notes before comparing results. A partial run's observed SR is not a full-suite score; custom harnesses, modified budgets, and alternative judges must be identified in submissions.

Ready to share? Follow the submission guide and open a result-submission GitHub issue. Include your model and harness versions, task revision, settings, trials, scores, and raw trajectories. Reviewers check the evidence before leaderboard publication.

Find your way around

agentworld/
├── agents/              Reference runner, model adapters, tools, and configs
├── benchmarks/          Offline scoring and benchmark entry-point docs
├── data_v0.1_multi/      Versioned multi-agent tasks and task verifiers
├── data_v0.1_solo/       Solo task definitions
├── examples/agents/     Custom-harness integration example
├── packages/            TypeScript game server, browser client, and shared code
├── analysis/            Research analysis and historical report archive
├── scripts/             Development utilities and manual diagnostics
├── docs/                Benchmark guides and game documentation
└── tests/benchmark/     Offline onboarding regression tests

The repository map distinguishes the supported onboarding path from historical experiments. Generated runs and visualizations belong in ignored runs/ directories. The leaderboard website lives in the separate agentworld-web repository.

Explore and contribute


Built upon Kaetram, which expands on Little Workshop's BrowserQuest. Code is licensed under MPL-2.0. Paper figures are by Shu et al., licensed under CC BY 4.0; see figure sources and attribution.

About

No description, website, or topics provided.

Resources

Contributing

Stars

27 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages