BridgeBench V3 Arena Refactoring
Code · 2026-07-10
The BridgeBench V3 Refactoring arena asks models to identify which candidate rewrite silently changes program behavior. Its 18 tasks span six refactoring clusters and use blind pairwise judging.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Fable 5 | 1144 |
| GLM-5.2 | 1003 |
| Grok 4.5 | 989 |
| GPT-5.6 Sol | 861 |