GPT-6 Sol

38.5± 0.7 (95% CI)
mini-SWE · pass rate 85.2% · simplicity 51.6 · $0.30 per library

Runs

Each cell is one full run: the model designs the library, then 3 implementers solve every problem with it. Click a cell to open that run.

TaskRun 1Run 2Run 3Mean ± std
npr38.541.545.842.0 ± 3.7
pyda43.844.445.744.6 ± 1.0
weft29.829.428.229.1 ± 0.8
sapi38.133.640.737.5 ± 3.6
simu41.238.539.939.9 ± 1.3
wapi26.623.228.025.9 ± 2.5