Can agents design libraries other agents can use?

One agent designs a library. Other agents solve problems with it. The library is scored on how correct and how short their code is.

Score against what it cost the model to design one library.

25303540455055
Production library · 46.6
No library · 34.4
Opus 5.5mini-SWE
Score ± 95% CI
48.9 ± 0.9
Pass rate
86.6%
Simplicity
64.5
Library $
$9.62
Fable 5.1mini-SWE
Score ± 95% CI
47.5 ± 1.1
Pass rate
86.1%
Simplicity
62.7
Library $
$14.81
GPT-6 Astramini-SWE
Score ± 95% CI
45.1 ± 0.7
Pass rate
85.7%
Simplicity
58.7
Library $
$3.63
GPT-6 AstraCodex
Score ± 95% CI
44.1 ± 0.8
Pass rate
85.5%
Simplicity
58.5
Library $
$4.51
Kimi K3mini-SWE
Score ± 95% CI
44.0 ± 1.0
Pass rate
86.0%
Simplicity
58.6
Library $
$8.70
GLM 5.3mini-SWE
Score ± 95% CI
41.9 ± 1.1
Pass rate
84.1%
Simplicity
57.1
Library $
$21.06
Fable 5.1Claude Code
Score ± 95% CI
39.9 ± 2.5
Pass rate
84.9%
Simplicity
58.3
Library $
$14.39
Grok 4.6mini-SWE
Score ± 95% CI
39.7 ± 1.1
Pass rate
84.5%
Simplicity
54.1
Library $
$2.63
GPT-5.6 SolCodex
Score ± 95% CI
39.5 ± 1.0
Pass rate
84.1%
Simplicity
52.9
Library $
$2.14
GPT-6 Solmini-SWE
Score ± 95% CI
38.5 ± 0.7
Pass rate
85.2%
Simplicity
51.6
Library $
$0.30
DeepSeek V4 Promini-SWE
Score ± 95% CI
31.2 ± 0.8
Pass rate
84.9%
Simplicity
42.0
Library $
$0.31
$0.2$0.5$1$2$5$10$20Library $ (log) →

Dashed line: best score for the cost. Hover a point for details.

Note: Scores run 0–100: pass rate² × simplicity, averaged over every problem and implementer, with tasks weighted equally. The error bars and ± show the 95% confidence interval.