Extended results
LibraryDesignBench, every metric. Best value in each column in bold.
#Model · click to see runsScore ± 95% CIPass rateSimplicityLibrary $Problem $
1
Opus 5.5mini-SWE
48.9 ±0.9
86.6% ±1.4
64.5 ±1.8
$9.62 ±1.03
$0.174 ±0.011
2
Fable 5.1mini-SWE
47.5 ±1.1
86.1% ±1.6
62.7 ±1.8
$14.81 ±1.37
$0.198 ±0.013
3
GPT-6 Astramini-SWE
45.1 ±0.7
85.7% ±1.7
58.7 ±1.6
$3.63 ±0.24
$0.155 ±0.008
4
GPT-6 AstraCodex
44.1 ±0.8
85.5% ±1.9
58.5 ±1.6
$4.51 ±0.40
$0.194 ±0.012
5
Kimi K3mini-SWE
44.0 ±1.0
86.0% ±1.5
58.6 ±1.7
$8.70 ±1.47
$0.176 ±0.010
6
GLM 5.3mini-SWE
41.9 ±1.1
84.1% ±1.7
57.1 ±1.7
$21.06 ±1.89
$0.234 ±0.013
7
Fable 5.1Claude Code
39.9 ±2.5
84.9% ±1.6
58.3 ±2.5
$14.39 ±2.23
$0.211 ±0.014
8
Grok 4.6mini-SWE
39.7 ±1.1
84.5% ±1.6
54.1 ±1.6
$2.63 ±0.23
$0.225 ±0.012
9
GPT-5.6 SolCodex
39.5 ±1.0
84.1% ±2.1
52.9 ±1.2
$2.14 ±0.20
$0.200 ±0.011
10
GPT-6 Solmini-SWE
38.5 ±0.7
85.2% ±1.8
51.6 ±1.3
$0.30 ±0.02
$0.150 ±0.007
11
DeepSeek V4 Promini-SWE
31.2 ±0.8
84.9% ±1.7
42.0 ±1.1
$0.31 ±0.02
$0.175 ±0.008
—
P
Production libraryReference arm
46.6 ±0.5
85.4% ±1.0
61.5 ±0.9
—
$0.239 ±0.021
—
–
No libraryReference arm
34.4 ±0.5
86.4% ±1.0
46.3 ±1.1
—
$0.098 ±0.006
Score ± is the 95% confidence interval; the other ± are standard errors. Score, pass rate and simplicity are percentages; pass rate and simplicity average finished cells, costs average every cell that recorded one.