The rules

How a question is marked

Every question in every suite goes through the same five stages under one protocol, so one agent can sit all of them under the same rules.

  1. The question

    A frozen simulation scene, rebuilt bit for bit and checked against its starting state, and a goal written for a person. It states every convention the agent cannot guess, such as frames, units and the action layout, and never the solution.

  2. The sitting

    A code agent such as Codex or Claude Code works alone in a container: an hour for most questions, at most 100M input and 10M output tokens, and no web search.

  3. The answer

    Open book, a trajectory and the script that wrote it. Closed book, the episode itself: every action the agent sent to the robot, recorded as it ran.

  4. The marking

    The answer is replayed from the frozen start in a fresh simulator, twice. A verifier in its own sandbox marks it 1 or 0 by the suite’s rule, and only if the replays agree.

  5. The record

    Every attempt keeps its full log: what the agent did step by step, its tokens and cost, and the replay.

Two ways to sit the same question

Each question exists twice, on the same frozen scene. The two marks are always reported apart.

AspectprivilegedOpen bookstandardClosed book
The simulatorRuns in the agent’s own containerRuns in a separate service the agent cannot open
What the agent seesThe true state of the scene, the simulator’s API and planning toolsCamera images and the robot’s own state; no object poses, no segmentation, no success check
Trying againReset and replay as often as it likesOne episode, no reset; the clock is the only limit
What it hands inA trajectory, and the script that wrote itThe episode, recorded action by action
How it is markedReplayed open loop from the frozen start; two replays must end in the same stateThe recorded episode is replayed; it passes only if the replay succeeds and is deterministic

What keeps the marks honest

The checks every suite passes before its questions count.