The rules
How a question is marked
Every question in every suite goes through the same five stages under one protocol, so one agent can sit all of them under the same rules.
- The question
A frozen simulation scene, rebuilt bit for bit and checked against its starting state, and a goal written for a person. It states every convention the agent cannot guess, such as frames, units and the action layout, and never the solution.
- The sitting
A code agent such as Codex or Claude Code works alone in a container: an hour for most questions, at most 100M input and 10M output tokens, and no web search.
- The answer
Open book, a trajectory and the script that wrote it. Closed book, the episode itself: every action the agent sent to the robot, recorded as it ran.
- The marking
The answer is replayed from the frozen start in a fresh simulator, twice. A verifier in its own sandbox marks it 1 or 0 by the suite’s rule, and only if the replays agree.
- The record
Every attempt keeps its full log: what the agent did step by step, its tokens and cost, and the replay.
Two ways to sit the same question
Each question exists twice, on the same frozen scene. The two marks are always reported apart.
| Aspect | privilegedOpen book | standardClosed book |
|---|---|---|
| The simulator | Runs in the agent’s own container | Runs in a separate service the agent cannot open |
| What the agent sees | The true state of the scene, the simulator’s API and planning tools | Camera images and the robot’s own state; no object poses, no segmentation, no success check |
| Trying again | Reset and replay as often as it likes | One episode, no reset; the clock is the only limit |
| What it hands in | A trajectory, and the script that wrote it | The episode, recorded action by action |
| How it is marked | Replayed open loop from the frozen start; two replays must end in the same state | The recorded episode is replayed; it passes only if the replay succeeds and is deterministic |
What keeps the marks honest
The checks every suite passes before its questions count.
- A separate marker
The verifier runs in its own sandbox, without network, and sees only what the agent declared as its answer.
- No answer key in reach
Reference solutions never enter the agent’s container, and upstream experts are removed from the images.
- Controls
Doing nothing must score 0, and a full reference solution must score 1.
- A cheat trial
Before a suite comes in, an agent is told to pass the verifier without doing the task. Any success blocks the suite until the verifier is fixed.
- Human review
A question without a full reference solution is checked by a reviewer who did not write it.
- Frozen questions
A published question never changes. Fixing a scene makes a new question.