Can agents design libraries other agents can use?
One agent designs a library. Other agents solve problems with it. The library is scored on how correct and how short their code is.
Score against what it cost the model to design one library.
25303540455055
Production library · 46.6
No library · 34.4
Opus 5.5mini-SWE
- Score ± 95% CI
- 48.9 ± 0.9
- Pass rate
- 86.6%
- Simplicity
- 64.5
- Library $
- $9.62
Fable 5.1mini-SWE
- Score ± 95% CI
- 47.5 ± 1.1
- Pass rate
- 86.1%
- Simplicity
- 62.7
- Library $
- $14.81
GPT-6 Astramini-SWE
- Score ± 95% CI
- 45.1 ± 0.7
- Pass rate
- 85.7%
- Simplicity
- 58.7
- Library $
- $3.63
GPT-6 AstraCodex
- Score ± 95% CI
- 44.1 ± 0.8
- Pass rate
- 85.5%
- Simplicity
- 58.5
- Library $
- $4.51
Kimi K3mini-SWE
- Score ± 95% CI
- 44.0 ± 1.0
- Pass rate
- 86.0%
- Simplicity
- 58.6
- Library $
- $8.70
GLM 5.3mini-SWE
- Score ± 95% CI
- 41.9 ± 1.1
- Pass rate
- 84.1%
- Simplicity
- 57.1
- Library $
- $21.06
Fable 5.1Claude Code
- Score ± 95% CI
- 39.9 ± 2.5
- Pass rate
- 84.9%
- Simplicity
- 58.3
- Library $
- $14.39
Grok 4.6mini-SWE
- Score ± 95% CI
- 39.7 ± 1.1
- Pass rate
- 84.5%
- Simplicity
- 54.1
- Library $
- $2.63
GPT-5.6 SolCodex
- Score ± 95% CI
- 39.5 ± 1.0
- Pass rate
- 84.1%
- Simplicity
- 52.9
- Library $
- $2.14
GPT-6 Solmini-SWE
- Score ± 95% CI
- 38.5 ± 0.7
- Pass rate
- 85.2%
- Simplicity
- 51.6
- Library $
- $0.30
DeepSeek V4 Promini-SWE
- Score ± 95% CI
- 31.2 ± 0.8
- Pass rate
- 84.9%
- Simplicity
- 42.0
- Library $
- $0.31
$0.2$0.5$1$2$5$10$20Library $ (log) →
Dashed line: best score for the cost. Hover a point for details.
Note: Scores run 0–100: pass rate² × simplicity, averaged over every problem and implementer, with tasks weighted equally. The error bars and ± show the 95% confidence interval.