The Cube Arena Index
One score for how well a model solves a cube it can only see as a picture: no code, no tools, just the model. Every result here was run by us or by people we invited, with our locked-down runner, the model's exact version and effort, the same rules and the same cubes. Missing a model? Help run it.
Running now: GPT-6-Luna (max) · Thorough, test 1 of 16
Cube Arena Index
Fast v1: Three ladders with 15 minutes each: 2×2 and 3×3 interactive, and 3×3 one-shot. About 45 minutes at most. Score out of 100, higher is better. Verified runs only.
- OpenAI
Score by test
Each test out of 100. Ladders: how far up the model climbed, with up to 20% for using few moves. Full scrambles: whether it solved the cube.
| Model | Index | 2×2 ladder, interactive | 3×3 ladder, interactive | 3×3 ladder, one-shot |
|---|---|---|---|---|
| GPT-6-Luna (max) | 17.3 | 27.3 | 9.5 | 15.0 |
What the score costs
Top left is where you want to be: a high score for few moves, little time or little money. The dotted line joins the models nobody beats on both.
Does more effort help?
The same model at each reasoning effort, low to max. A flat line means thinking longer doesn't help it see the cube.
All results
Every Verified Fast result shown above. A model and effort run more than once shows the average. Click a score for each test and its replays, or a model for all its efforts.
| # | Model | Maker | Index | Solved | Moves | Time | Tokens | Cost | Runs | Run by |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | GPT-6-Luna (max) | OpenAI | 17.3 | 8 | 232 | 43 min | 4.3M | $0.19* | 1 | Cube Arena · via Codex |
How the index works
Two suites
Fast (45 minutes at most): 3 ladders, 2×2 and 3×3 interactive and 3×3 one-shot, 15 minutes each. Every model gets it, at every effort level.
Thorough (up to 8 hours): all 16 model-only tests, 2×2 to 5×5, ladder and full scramble, interactive and one-shot. Ladders get 30 minutes; each full scramble is two cubes of 15 minutes. It goes to the strongest models from Fast.
Grading
Each test scores 0 to 100. A ladder scores how far up it climbed: level 7 of 20 is 35. A full scramble scores 100 for a solve, 0 for a miss.
Up to 20% of each test is for few moves: a level-N cube needs N moves, a 3×3 scramble 20. A wasteful solve keeps at least 80%. The index is the average of the tests.
Time counts through each test's budget: when it runs out, the test ends where the model got to.
Fair runs
- Same cubes for every model: fixed per suite version, from a secret only the server knows.
- The official agent CLI (Claude Code or Codex) in an empty folder, with the cube tools as its only tools. Anything else stops the run.
- A fresh session for every test, so long suites never run out of context.
- Exact names from the model id and effort. Claude runs check which model actually answered.
Reading the numbers
- Charts show Verified runs only. Community runs stay on the boards.
- Cost: Claude Code's API-equivalent cost, or GPT tokens at OpenAI's list prices. The runs used subscriptions.
- Changing tests or grading makes a new version; scores only compare within one.
- Every cube has a replay. Open any result, or run your own model.