Can agents design libraries other agents can use?

One agent designs a library. Other agents solve problems with it. The library is scored on how correct and how short their code is.

Score against what the implementers spent per problem using each model's libraries, with the no-library and production-library arms for comparison.

25303540455055
Opus 5.5mini-SWE
Score ± 95% CI
48.9 ± 0.9
Pass rate
86.6%
Simplicity
64.5
$ / problem
$0.174
Fable 5.1mini-SWE
Score ± 95% CI
47.5 ± 1.1
Pass rate
86.1%
Simplicity
62.7
$ / problem
$0.198
GPT-6 Astramini-SWE
Score ± 95% CI
45.1 ± 0.7
Pass rate
85.7%
Simplicity
58.7
$ / problem
$0.155
GPT-6 AstraCodex
Score ± 95% CI
44.1 ± 0.8
Pass rate
85.5%
Simplicity
58.5
$ / problem
$0.194
Kimi K3mini-SWE
Score ± 95% CI
44.0 ± 1.0
Pass rate
86.0%
Simplicity
58.6
$ / problem
$0.176
GLM 5.3mini-SWE
Score ± 95% CI
41.9 ± 1.1
Pass rate
84.1%
Simplicity
57.1
$ / problem
$0.234
Fable 5.1Claude Code
Score ± 95% CI
39.9 ± 2.5
Pass rate
84.9%
Simplicity
58.3
$ / problem
$0.211
Grok 4.6mini-SWE
Score ± 95% CI
39.7 ± 1.1
Pass rate
84.5%
Simplicity
54.1
$ / problem
$0.225
GPT-5.6 SolCodex
Score ± 95% CI
39.5 ± 1.0
Pass rate
84.1%
Simplicity
52.9
$ / problem
$0.200
GPT-6 Solmini-SWE
Score ± 95% CI
38.5 ± 0.7
Pass rate
85.2%
Simplicity
51.6
$ / problem
$0.150
DeepSeek V4 Promini-SWE
Score ± 95% CI
31.2 ± 0.8
Pass rate
84.9%
Simplicity
42.0
$ / problem
$0.175
P
P
Production library
Score ± 95% CI
46.6 ± 0.5
Pass rate
85.4%
Simplicity
61.5
$ / problem
$0.239
–
–
No library
Score ± 95% CI
34.4 ± 0.5
Pass rate
86.4%
Simplicity
46.3
$ / problem
$0.098
$0.1$0.2$ / problem (log) →

Dashed line: best score for the cost. Hover a point for details.

Note: Scores run 0–100: pass rate² × simplicity, averaged over every problem and implementer, with tasks weighted equally. The error bars and ± show the 95% confidence interval.