Atlas

Benchmarks

← All benchmarks

LibraryDesignBench

Code

Fifteen open-ended library design tasks in Rust, Python, TypeScript and Haskell, evaluated through 242 downstream problems solved by three fixed implementer agents. Score is pass rate squared times simplicity, averaged across problems and implementers with equal task weights. Designers have four hours per library offline; implementers have one hour and $2.50 per problem. Reported uncertainty is a 95% confidence interval.

Top models (higher is better)

ModelScore
Claude Opus 5.548.9
Claude Fable 5.147.5
GPT-6 Astra45.1
Kimi K344.0
GLM-5.341.9
Grok 4.639.7
GPT-5.6 Sol39.5
GPT-6 Sol38.5
Loading Atlas data…