Scoring
score = pass rate² × simplicity
- Pass rate
- Behavioral tests passed / total.
- Simplicity
- The mean of four ratios of reference to solution: source lines, cyclomatic complexity, cognitive complexity and Halstead volume. Each ratio is capped at 1, so a solution earns full credit for matching the reference's size and none for undercutting it.
- Reference
- The problem's production-library solution, measured after the same formatter.
From solutions to models
- A library's score is the mean over all of its problem × implementer cells. A cell that did not finish scores 0.
- A model's score averages its libraries within each task, then weights every task equally.
- The same implementers also solve every problem with no library and with the production library, which anchor the comparison.