LibraryDesignBench
Code
Fifteen open-ended library design tasks in Rust, Python, TypeScript and Haskell, evaluated through 242 downstream problems solved by three fixed implementer agents. Score is pass rate squared times simplicity, averaged across problems and implementers with equal task weights. Designers have four hours per library offline; implementers have one hour and $2.50 per problem. Reported uncertainty is a 95% confidence interval.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Opus 5.5 | 48.9 |
| Claude Fable 5.1 | 47.5 |
| GPT-6 Astra | 45.1 |
| Kimi K3 | 44.0 |
| GLM-5.3 | 41.9 |
| Grok 4.6 | 39.7 |
| GPT-5.6 Sol | 39.5 |
| GPT-6 Sol | 38.5 |