Atlas

Benchmarks

← All benchmarks

BridgeBench V3 Arena Refactoring

Code · 2026-07-10

The BridgeBench V3 Refactoring arena asks models to identify which candidate rewrite silently changes program behavior. Its 18 tasks span six refactoring clusters and use blind pairwise judging.

Top models (higher is better)

ModelScore
Claude Fable 51144
GLM-5.21003
Grok 4.5989
GPT-5.6 Sol861
Loading Atlas data…