Can coding agents finish Paper Mario?
No, but some of them can finish part of it.
The benchmark is the game's opening: from a fresh save file, through Peach's castle, until the first Bowser battle begins. A human clears it in minutes; agents get 90. And this is not a computer-use eval: the agents play with a full toolkit, from an interactive controller and TAS-style controls to the game's decompiled source and walkthroughs.

- G4 finishes
- 119
- Model families
- 23
- Run cap
- 90m
The standings
How far each model gets, and how fast it gets there.
The score is route progress: reach the first Bowser battle and that's 100%, with partial credit for every scored step along the way. The rankings weigh three things. Can a setup finish at all, does it do so consistently, and how fast does it get there. Each row shows average progress, a marker for the best run, and a separate time bar for the fastest finish. The clock settles ties between setups that reach equally far.
| Model | Progress & completion time | Avg | Best | G4 | Best G4 time | Furthest | Est. cost |
|---|---|---|---|---|---|---|---|
| 100% | 100% | 3/3 | 20:05 | G4 | ~$12.29Mean estimated cost per attempt in this effort. Conservative estimate: all recorded non-cached input is priced as cache writes because Codex did not reliably report the split. The recorded-usage range is $11.99 to $12.29. | ||
| 100% | 100% | 3/3 | 21:58 | G4 | $5.44Mean estimated cost per attempt in this effort. API-equivalent estimate from recorded usage. Unreported usage is excluded; this is not a bill. | ||
| 100% | 100% | 3/3 | 24:10 | G4 | $7.61Mean estimated cost per attempt in this effort. API-equivalent estimate from recorded usage. Unreported usage is excluded; this is not a bill. | ||
| 100% | 100% | 3/3 | 24:36 | G4 | $13.90Mean estimated cost per attempt in this effort. API-equivalent estimate from recorded usage. Unreported usage is excluded; this is not a bill. | ||
| 100% | 100% | 3/3 | 26:32 | G4 | $17.05Mean estimated cost per attempt in this effort. API-equivalent estimate from recorded usage. Unreported usage is excluded; this is not a bill. | ||
| 100% | 100% | 3/3 | 36:41 | G4 | $10.32Mean estimated cost per attempt in this effort. API-equivalent estimate from recorded usage. Unreported usage is excluded; this is not a bill. | ||
| 100% | 100% | 3/3 | 41:39 | G4 | $15.60Mean estimated cost per attempt in this effort. API-equivalent estimate from recorded usage. Unreported usage is excluded; this is not a bill. | ||
| 100% | 100% | 3/3 | 44:35 | G4 | $0.76Mean estimated cost per attempt in this effort. API-equivalent estimate from recorded usage. Unreported usage is excluded; this is not a bill. | ||
| 96% | 100% | 2/3 | 1:09:32 | G4 | $14.14Mean estimated cost per attempt in this effort. API-equivalent estimate from recorded usage. Unreported usage is excluded; this is not a bill. | ||
| 93% | 100% | 2/3 | 36:23 | G4 | $6.43Mean estimated cost per attempt in this effort. API-equivalent estimate from recorded usage. Unreported usage is excluded; this is not a bill. | ||
| 89% | 100% | 2/3 | 44:14 | G4 | $6.18Mean estimated cost per attempt in this effort. API-equivalent estimate from recorded usage. Unreported usage is excluded; this is not a bill. | ||
| 89% | 100% | 2/3 | 1:06:46 | G4 | $10.57Mean estimated cost per attempt in this effort. API-equivalent estimate from recorded usage. Unreported usage is excluded; this is not a bill. | ||
| 85% | 100% | 2/3 | 47:33 | G4 | $0.64Mean estimated cost per attempt in this effort. API-equivalent estimate from recorded usage. Unreported usage is excluded; this is not a bill. | ||
| 85% | 100% | 1/3 | 1:09:10 | G4 | $6.18Mean estimated cost per attempt in this effort. API-equivalent estimate from recorded usage. Unreported usage is excluded; this is not a bill. | ||
| 82% | 100% | 2/3 | 44:19 | G4 | $10.38Mean estimated cost per attempt in this effort. API-equivalent estimate from recorded usage. Unreported usage is excluded; this is not a bill. | ||
| 82% | 100% | 1/3 | 1:08:40 | G4 | $51.41Mean estimated cost per attempt in this effort. API-equivalent estimate from recorded usage. Unreported usage is excluded; this is not a bill. | ||
| 67% | 100% | 1/3 | 47:53 | G4 | $9.04Mean estimated cost per attempt in this effort. API-equivalent estimate from recorded usage. Unreported usage is excluded; this is not a bill. | ||
No finish | 63% | 78% | 0/3 | — | S3 | $5.47Mean estimated cost per attempt in this effort. API-equivalent estimate from recorded usage. Unreported usage is excluded; this is not a bill. | |
No finish | 59% | 78% | 0/3 | — | S3 | $1.86Mean estimated cost per attempt in this effort. API-equivalent estimate from recorded usage. Unreported usage is excluded; this is not a bill. | |
No finish | 52% | 78% | 0/3 | — | S3 | $0.87Mean estimated cost per attempt in this effort. API-equivalent estimate from recorded usage. Unreported usage is excluded; this is not a bill. | |
No finish | 52% | 78% | 0/3 | — | S3 | $12.16Mean estimated cost per attempt in this effort. API-equivalent estimate from recorded usage. Unreported usage is excluded; this is not a bill. | |
No finish | 52% | 89% | 0/3 | — | S4 | $5.33Mean estimated cost per attempt in this effort. API-equivalent estimate from recorded usage. Unreported usage is excluded; this is not a bill. | |
No finish | 44% | 44% | 0/3 | — | S1 | $10.47Mean estimated cost per attempt in this effort. API-equivalent estimate from recorded usage. Unreported usage is excluded; this is not a bill. |
*GLM 5.2 and Kimi K2.7 Code were given custom image MCP tools backed by locally hosted Qwen3.6-35B and LocateAnything-3B.
Muse Spark is omitted from these rankings because its 50-image-per-request cap prevented comparable testing. Which versions we tried and what happened.
How the cost estimates are counted
Cost estimates cover 255 of 259 retained runs. They use recorded usage and run-era pricing or client price tables; missing usage can leave an estimate incomplete. Missing costs stay blank.
These are estimates, not bills. OpenRouter uses an unversioned client price table. Astra’s ~ estimate prices all recorded non-cached input as cache writes because the input/write split was not reliably reported. Read the cost breakdown and its limits.
The route, in three frames
Save file to Bowser: the ground every run has to cover.



Rendered with the Paper Mario 64 R HD texture pack by MasterKillua; the pack itself is not redistributed here.
The common clock
Same route: 3:47.4 for a lifelong player, 20:05.3 for the fastest agent.
One route-familiar human on direct controller input, with a casual player as a second reference point. Agents play through a paused harness that freezes the game while they think, so the gap is execution overhead, not a ranking of capability. Human rows are never averaged with each other or with model cells.
Explore the August race by checkpoint
Shared scorer clock · same horizontal scale
One clock. 4 finishes.
Under the hood
The deep dive
Explore individual runs, effort settings, and the six-hour marathons beyond the opening.
The tools
See how agents observe and advance the game, then explore the six harnesses and the rules they share.
The behavior
Follow what agents notice, what they try, and how they recover when the game stops going their way.