Guide
A task
An open spec for a library, a few example usages, and a set of downstream problems. Each problem has behavioral tests and reference solutions written with a real production library and with no library.
clirs/
problem.yaml # language, production library, problems
phase_1/instruction.md # the open spec the designer gets
phase_2/site-builder/
instruction.md # what the implementer gets
tests/ # behavioral tests + static_reference.json
solution/ # reference arms: clap, no-library
workspace/ # neutral starter codeDesign phase
The model under test writes and packages the library from the open spec. It gets capabilities and example usages, not interfaces, and never sees the downstream problems or their tests. Each model designs three libraries per task.
Use phase
Three fixed implementers solve every problem with each library: DeepSeek V4.1 Flash (mini-SWE), GLM 5.3 Flash (mini-SWE), GPT-5.6 Luna (Codex). The same implementers also solve every problem with no library and with the task's production library, which anchor the comparison.
Score
Each solution scores pass rate² × simplicity. Simplicity is the mean of four ratios of reference to solution: source lines, cyclomatic complexity, cognitive complexity and Halstead volume. Each ratio is capped at 1, so a solution earns full credit for matching the reference's size and none for undercutting it.
A library's score is the mean over all of its problem × implementer cells, and a cell that did not finish scores 0. A model's score averages its libraries within each task, then weights every task equally. ± is the standard error of re-running the benchmark on the same tasks.
LibraryUseBench
No design phase. Each model is the implementer: it gets the task's production library and a minimal instruction to use it and write as little code as possible, and is scored on the same problems with the same formula.
Run
git clone https://github.com/SprocketLab/librarydesignbench cd librarydesignbench && uv sync uv run ldb run <config | experiment> # design libraries, then evaluate them uv run ldb eval no-library # the no-library floor uv run ldb eval existing-library # the production-library ceiling uv run ldb static <run> # refresh static evidence and scores