BridgeBench V3 Arena BullShit
Safety · 2026-07-10
The BridgeBench V3 BullShit arena tests whether models complete valid parts of a request while challenging fabricated concepts, impossible quantities, reversed causality, pseudoscience, and loaded assumptions. Its 18 tasks span six clusters and use blind pairwise judging.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Fable 5 | 1154 |
| GLM-5.2 | 1003 |
| Grok 4.5 | 978 |
| GPT-5.6 Sol | 868 |