Atlas

Benchmarks

← All benchmarks

LibraryDesignBench Simplicity

Code

Task-weighted simplicity of downstream code written using the designed library, measured against reference solutions using production libraries. The benchmark evaluates downstream implementation simplicity rather than library code size.

Top models (higher is better)

ModelScore
Claude Opus 5.564.5
Claude Fable 5.162.7
GPT-6 Astra58.7
Kimi K358.6
GLM-5.357.1
Grok 4.654.1
GPT-5.6 Sol52.9
GPT-6 Sol51.6
Loading Atlas data…