GPT-5.6 Sol

39.5± 1.0 (95% CI)
Codex · pass rate 84.1% · simplicity 52.9 · $2.14 per library

Runs

Each cell is one full run: the model designs the library, then 3 implementers solve every problem with it. Click a cell to open that run.

TaskRun 1Run 2Run 3Mean ± std
npr47.040.542.843.4 ± 3.3
pyda45.040.944.543.5 ± 2.2
weft33.231.436.533.7 ± 2.6
sapi36.838.833.436.4 ± 2.7
simu40.643.844.342.9 ± 2.0
wapi13.721.710.515.3 ± 5.7