No design phase. Each agent gets a real production library and one instruction: use it, and write as little code as possible.
Each cell is the row model's score minus the column model's, in points. Both solve the same 242 problems, so the difference is taken problem by problem and the difficulty they share cancels. That makes these intervals much narrower than the leaderboard's: models whose bars overlap there can still be clearly apart here.
| Row − column | Opus 5.5 | Sonnet 5.5 | GPT-6 Sol | Opus 5 | DeepSeek V4.1 Flash | GPT-5.6 Terra | Sonnet 5 | GLM 5.3 Flash | DeepSeek V4 Flash | GPT-6 Luna | GPT-5.6 Luna |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Opus 5.5 | +3.1±2.1 | +9.4±3.5 | +11.9±5.0 | +23.1±4.9 | +26.9±5.3 | +27.1±6.2 | +31.9±4.7 | +32.2±5.8 | +33.4±7.1 | +36.8±6.0 | |
| Sonnet 5.5 | −3.1±2.1 | +6.2±3.1 | +8.7±5.1 | +20.0±3.9 | +23.8±4.6 | +23.9±5.5 | +28.8±4.0 | +29.0±5.4 | +30.3±6.3 | +33.7±5.2 | |
| GPT-6 Sol | −9.4±3.5 | −6.2±3.1 | +2.5±4.3 | +13.8±3.4 | +17.6±2.9 | +17.7±4.9 | +22.6±2.7 | +22.8±5.2 | +24.1±4.7 | +27.4±3.1 | |
| Opus 5 | −11.9±5.0 | −8.7±5.1 | −2.5±4.3 | +11.3±4.5 | +15.1±4.3 | +15.2±3.4 | +20.1±4.3 | +20.3±5.8 | +21.6±6.1 | +24.9±4.6 | |
| DeepSeek V4.1 Flash | −23.1±4.9 | −20.0±3.9 | −13.8±3.4 | −11.3±4.5 | +3.8±3.6 | +3.9±3.9 | +8.8±2.0 | +9.0±3.4 | +10.3±4.1 | +13.7±3.0 | |
| GPT-5.6 Terra | −26.9±5.3 | −23.8±4.6 | −17.6±2.9 | −15.1±4.3 | −3.8±3.6 | +0.1±3.2 | +5.0±3.1 | +5.2±5.1 | +6.5±3.9 | +9.9±2.3 | |
| Sonnet 5 | −27.1±6.2 | −23.9±5.5 | −17.7±4.9 | −15.2±3.4 | −3.9±3.9 | −0.1±3.2 | +4.9±3.8 | +5.1±5.4 | +6.4±4.7 | +9.8±3.8 | |
| GLM 5.3 Flash | −31.9±4.7 | −28.8±4.0 | −22.6±2.7 | −20.1±4.3 | −8.8±2.0 | −5.0±3.1 | −4.9±3.8 | +0.2±4.1 | +1.5±3.7 | +4.9±2.3 | |
| DeepSeek V4 Flash | −32.2±5.8 | −29.0±5.4 | −22.8±5.2 | −20.3±5.8 | −9.0±3.4 | −5.2±5.1 | −5.1±5.4 | −0.2±4.1 | +1.3±5.4 | +4.6±5.1 | |
| GPT-6 Luna | −33.4±7.1 | −30.3±6.3 | −24.1±4.7 | −21.6±6.1 | −10.3±4.1 | −6.5±3.9 | −6.4±4.7 | −1.5±3.7 | −1.3±5.4 | +3.4±3.1 | |
| GPT-5.6 Luna | −36.8±6.0 | −33.7±5.2 | −27.4±3.1 | −24.9±4.6 | −13.7±3.0 | −9.9±2.3 | −9.8±3.8 | −4.9±2.3 | −4.6±5.1 | −3.4±3.1 |
Green: the row model is ahead, and the 95% interval of the difference excludes zero. Red: it is behind. Plain: the two cannot be told apart. ± is 1.96 × the task-clustered standard error of the per-problem differences (how it is computed). Hover a cell for its interval.
How are the confidence intervals computed?
A model's score is its mean over all 242 problems. Each problem counts once: its three runs are averaged first, and a run that did not finish scores 0.
The problems of one task all use the same production library, so a model that handles that library well tends to do well on all of them. Their scores are not independent, and treating them as if they were would make the intervals look far narrower than the data supports. The standard error is therefore clustered by task (15 clusters), following Miller (2024), "Adding Error Bars to Evals":
where G is the number of tasks and n the number of problems. The ± shown is 1.96 × SE. The interval says how much the score could move on other tasks like these, not only on reruns of the same ones.
Paired differences use the same clustering on each problem's difference between two models. Both models solve the same problems, so the difficulty they share cancels, and a difference can be clear even when the two leaderboard intervals overlap.