Atlas

Benchmarks

← All benchmarks

BridgeBench V3 Arena Generation

Code · 2026-07-10

The BridgeBench V3 Generation arena asks models to choose the implementation that fully satisfies a specification, including edge cases and interacting constraints. Its 18 tasks span six code-generation clusters and use blind pairwise judging.

Top models (higher is better)

ModelScore
Claude Fable 51101
GLM-5.2990
Grok 4.5975
GPT-5.6 Sol958
Loading Atlas data…