How a run works

One run designs three libraries per task, then every implementer solves every problem with each of them.

How an experiment runs on clirs: the author agent writes three crates, each passing an offline readiness check; 3 crates × 3 implementors × 13 problems fan out into 117 trials; each trial materializes the problem, builds the image, installs the library, solves, verifies and scores pass rate² × simplicity.
  1. 1

    Build the image

    environment/Dockerfile is built once per task and shared by both phases.

    Pre-installed
    Agent tools, the language toolchain, and the task's pinned dependencies, fetched ahead so everything works offline.
    Left out
    The production library and anything like it. Only the production-library arm adds it, on top of this image.
  2. 2

    Design

    The model under test writes a library from the open spec. Three libraries per model per task.

    Sees
    design/instruction.md and an empty /workspace
    Hidden
    The problems, their tests, and the production library
    Limits
    4 hours · network to the model API only
    Gate
    The package must install offline, or the library is rejected
  3. 3

    Install the library

    Before each implementer starts, in a fresh container.

    /library
    The designer's /workspace, mounted read-only
    /workspace
    The problem's starter project, with the library installed as a dependency
    If it fails
    The solution scores 0
  4. 4

    Implement

    Three fixed implementers solve every problem with the library, one container each.

    Sees
    The problem's instruction.md, the starter project, and the library source
    Told
    The library is installed, and the solution is judged on how little code sits on top of it
    Hidden
    Tests, reference solutions, other problems
    Limits
    1 hour · package registries blocked
  5. 5

    Grade

    After the implementer exits, the tests are copied in.

    Format
    The pinned formatter (ruff, rustfmt, fourmolu, prettier) normalizes the code
    Measure
    tree-sitter counts lines and complexity of the implementer's code only
    Test
    Behavioral tests run the program; pass rate = passed / total
  6. 6

    Score

    pass rate² × simplicity, then averaged up to the library and the model.

Implementers: DeepSeek V4.1 Flash (mini-SWE), GLM 5.3 Flash (mini-SWE), GPT-5.6 Luna (Codex).

Library conditions

Every arm runs in the same image. Only /library and the project's dependencies differ.

Arm/libraryInstalled into the project
Designed libraryThe designer's package, read-onlyYes, before the agent starts
No libraryEmptyNothing
Production librarySource of the pinned releaseYes, at image build (e.g. clap 4.6.1)

Before every no-library and designed-library solution starts, a check fails the run if the task's production library is importable.