Atlas

Benchmarks

← All benchmarks

BridgeBench V3 Arena Debugging

Code · 2026-07-10

The BridgeBench V3 Debugging arena asks models to isolate the true cause of a failing system despite planted decoys and symptom-hiding fixes. Its 18 tasks span six debugging clusters and use blind pairwise judging.

Top models (higher is better)

ModelScore
Claude Fable 51156
GLM-5.21002
Grok 4.5992
GPT-5.6 Sol848
Loading Atlas data…