No design phase. Each agent gets a real production library and one instruction: use it, and write as little code as possible.
25303540455055606570
Opus 5.5mini-SWE
- Score ± 95% CI
- 66.9 ± 7.1
- Pass rate
- 93.0%
- Simplicity
- 78.3
- $ / problem
- $0.664
Opus 5mini-SWE
- Score ± 95% CI
- 55.1 ± 6.7
- Pass rate
- 87.0%
- Simplicity
- 68.5
- $ / problem
- $1.314
DeepSeek V4.1 Flashmini-SWE
- Score ± 95% CI
- 43.8 ± 5.3
- Pass rate
- 90.5%
- Simplicity
- 52.1
- $ / problem
- $0.190
GPT-5.6 Terramini-SWE
- Score ± 95% CI
- 40.0 ± 6.6
- Pass rate
- 82.2%
- Simplicity
- 57.7
- $ / problem
- $0.239
Sonnet 5mini-SWE
- Score ± 95% CI
- 39.9 ± 6.8
- Pass rate
- 80.1%
- Simplicity
- 56.0
- $ / problem
- $0.919
GLM 5.3 Flashmini-SWE
- Score ± 95% CI
- 35.0 ± 5.2
- Pass rate
- 87.3%
- Simplicity
- 47.2
- $ / problem
- $0.192
DeepSeek V4 Flashmini-SWE
- Score ± 95% CI
- 34.8 ± 6.4
- Pass rate
- 88.0%
- Simplicity
- 49.3
- $ / problem
- $0.286
GPT-5.6 Lunamini-SWE
- Score ± 95% CI
- 30.1 ± 5.8
- Pass rate
- 79.0%
- Simplicity
- 46.5
- $ / problem
- $0.042
Muse Spark 1.3mini-SWE
- Score ± 95% CI
- 29.4 ± 5.6
- Pass rate
- 92.0%
- Simplicity
- 35.8
- $ / problem
- $1.940
$0.05$0.1$0.2$0.5$1$2$ / problem (log) →
Dashed line: best score for the cost. Hover a point for details.
Note: Scores run 0–100: pass rate² × simplicity, averaged over every problem, with problems weighted equally. The error bars and ± show the 95% confidence interval.