Can agents design libraries other agents can use?

Agents will soon replace humans as both the main consumers and designers of libraries. So a library's quality should be measured by its utility to agents, not humans. We built a benchmark to test exactly this: One agent designs a library, and we score it only by how well other agents build with it.

Score against what the implementers spent per problem using each model's libraries, with the no-library and production-library arms for comparison.

25303540455055
$0.1$0.15$0.2$0.3$ / problem (log) →

Line: best score for the cost. Hover or tap a point for details.

Scores run 0–100: pass rate² × simplicity, averaged over every problem and implementer, with tasks weighted equally.

About LibraryDesignBench

Each task asks the model under test to design a reusable library from an open-ended spec: the capabilities it should have and a few example uses. It gets no interfaces, no method signatures and no tests. There are 15 tasks in Rust, Python, TypeScript and Haskell, from a CLI argument parser like clap to a dataframe library like pandas.

The library itself is never scored: not its code, its size or its own tests. It only has to build. It is judged only by the code other agents write with it. Three fixed implementer agents solve each of the task's problems with the library, without ever seeing the tests. Each solution is scored on how simple it is compared with a reference solution built on the real production library, weighted by how many hidden tests it passes.

So a library scores well only when it makes the code built on it short and simple. The docs cover every step, including exactly what is pre-installed for each agent.

How it works
Diagram of one LibraryDesignBench task, from the spec the designer gets to one solution's score. Each part is described below.

Top left: the model under test gets only instruction.md, an open spec for a library (here clirs, a Rust CLI parser). It never sees tests, interfaces or an API design, and writes the whole library itself.

Top right: three downstream agents each solve the task's 13 problems using that library. Each square is one solution's score, darker is better. The library score is the mean of every square.

Bottom: one square, zoomed in. A solution scores its pass rate squared times its simplicity: four size and complexity ratios against a reference solution written with the real production library (here clap), each capped at 1. Less code scores higher.

The numbers are from one example run. The leaderboard uses the same scale, written as 0–100 points.

What are the "No library" and "Production library" rows?

The same three implementers solve the same problems with no library at all, or with the real human-written library the task is modeled on (clap, pandas, …). They show whether a designed library beats having nothing, and how it compares with what an expert would reach for.

Can a library make things worse?

Yes. DeepSeek V4 Pro's libraries score 9.2% below having no library at all. Harm is most common in Haskell, where an agent-written library scores below no library on 70% of designer and task pairs.

Why is the harness listed separately?

It changes the score. Fable 5.1 scores 47.5 in mini-SWE-agent but 39.9 in Claude Code, so each model and harness pair is ranked as its own entry. Designers run in mini-SWE-agent unless another harness is named.

Do implementers favor libraries from their own model family?

No. All three implementers agree on the top three designers and the last one, and none ranks its own family higher. Their absolute scores differ, and the leaderboard weights them equally.

What limits do the agents run under?

No internet access. Designers get 4 hours per library. Implementers get 1 hour and $2.50 per problem. Unfinished solutions are scored as they are, which affects 1.7% of runs.

Do agents just copy existing libraries?

Mostly. On 11 of 15 tasks, designers reproduce the production library's design. In the CLI parser task, for example, they copy clap's builder style rather than its shorter derive macro.

Why do agents still write more code than the reference?

Mostly because the library is rigid or hard to use, not because a feature is missing. In an audit of solutions longer than their reference, 64% of the excess code traced to rigid or verbose interfaces and 14% to missing capabilities.

What doesn't the score measure?

The library's own correctness, security, runtime speed or maintainability. It measures only how correct and simple the programs built on it are, for these tasks, implementers and budgets.

Citation

If you use LibraryDesignBench, please cite the paper:

@misc{orlanski2026agentsdesignlibrariesagents,
      title={Can Agents Design Libraries for Agents?},
      author={Gabriel Orlanski and Alex L. Zhang and Avi Trost and Vincent Sunn Chen and Frederic Sala and Aws Albarghouthi and Ludwig Schmidt},
      year={2026},
      eprint={2609.36730},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2609.36730},
}

Supported by

DARPANational Science FoundationSnorkel AIPrime Intellect

Special thanks to Snorkel AI for supporting this work through the Open Benchmarks Grant.