No design phase. Each agent gets a real production library and one instruction: use it, and write as little code as possible.

25303540455055606570
Opus 5.5mini-SWE
Score ± 95% CI
66.9 ± 7.1
Pass rate
93.0%
Simplicity
78.3
$ / problem
$0.664
Opus 5mini-SWE
Score ± 95% CI
55.1 ± 6.7
Pass rate
87.0%
Simplicity
68.5
$ / problem
$1.314
DeepSeek V4.1 Flashmini-SWE
Score ± 95% CI
43.8 ± 5.3
Pass rate
90.5%
Simplicity
52.1
$ / problem
$0.190
GPT-5.6 Terramini-SWE
Score ± 95% CI
40.0 ± 6.6
Pass rate
82.2%
Simplicity
57.7
$ / problem
$0.239
Sonnet 5mini-SWE
Score ± 95% CI
39.9 ± 6.8
Pass rate
80.1%
Simplicity
56.0
$ / problem
$0.919
GLM 5.3 Flashmini-SWE
Score ± 95% CI
35.0 ± 5.2
Pass rate
87.3%
Simplicity
47.2
$ / problem
$0.192
DeepSeek V4 Flashmini-SWE
Score ± 95% CI
34.8 ± 6.4
Pass rate
88.0%
Simplicity
49.3
$ / problem
$0.286
GPT-5.6 Lunamini-SWE
Score ± 95% CI
30.1 ± 5.8
Pass rate
79.0%
Simplicity
46.5
$ / problem
$0.042
Muse Spark 1.3mini-SWE
Score ± 95% CI
29.4 ± 5.6
Pass rate
92.0%
Simplicity
35.8
$ / problem
$1.940
$0.05$0.1$0.2$0.5$1$2$ / problem (log) →

Dashed line: best score for the cost. Hover a point for details.

Note: Scores run 0–100: pass rate² × simplicity, averaged over every problem, with problems weighted equally. The error bars and ± show the 95% confidence interval.