Contribute

Add to the exam

Bring the benchmark you know best. Every accepted suite is credited to the people who brought it.

Five ways in

From a whole suite to a single bug report.

What a suite needs

Six things every proposal shows.

From proposal to leaderboard

Six steps. The protocol ↗ has every detail.

  1. ProposeOpen an issue, or fill in the form below.proposal-approved
  2. BuildTwo task folders per frozen scene, from a generator.-privileged-standard
  3. Self-checkRun the checks the reviewers will run.run_all.shoracle = 1nop = 02 replays
  4. ReviewA maintainer, a cheat trial and a model trial.cheat trialhuman review
  5. MergeBase images published, the suite added to the site.eai-<name>
  6. On the boardIts questions join the exam’s evaluations.both modesevery log

Propose a suite

Fill in the fields of a proposal and copy it as Markdown. Nothing is sent from this page.

Where a new suite helps most

Labels and robot bodies with room for more questions.

Credit

How contributions could count toward authorship of the exam’s paper.

House rules

From the contributing guide and the protocol.

Questions

How do I send a proposal?

Fill in the form above and copy it as Markdown, then send it to a maintainer you know. The exam’s repository is not public yet; access for the pull request is arranged once a proposal is approved.

My benchmark has no scripted expert. Can it still come in?

Yes. A family without a full reference solution needs a positive example (a successful model trial, an upstream policy rollout or a hand-built trajectory) and a human review. Its results are reported apart from the tasks with a verified solution until it has one.

Does my suite need both modes?

Each frozen scene is offered as a privileged and a standard task. Standard mode needs a service that runs the simulator behind the standard protocol (eai-standard/2.x); the services of the suites already in are the templates.

Which simulators are accepted?

Any that runs headless in a pinned Docker image and replays a frozen scene deterministically on the reference hardware, an RTX 4090. The exam already runs MuJoCo, Isaac Sim, SAPIEN and Genesis, among others.

What if the simulator or its assets cannot be redistributed?

The image is marked build-only and comes with build instructions (licences, data). It is never published, though the maintainers may keep a private package for the team.

Can I get my own agent on the leaderboard?

Official results come from Harbor’s native agents (Codex, Claude Code and others) under the protocol’s budgets, run with our runner. Helpers with GPU hosts run assigned batches and send the job directories back; ask the maintainers for an assignment.

How is a suite credited?

Every task names its author, and every adapted suite credits its upstream benchmark on its suite page and in its documentation.