Atlas

Benchmarks

← All benchmarks

BridgeBench V3 Arena BullShit

Safety · 2026-07-10

The BridgeBench V3 BullShit arena tests whether models complete valid parts of a request while challenging fabricated concepts, impossible quantities, reversed causality, pseudoscience, and loaded assumptions. Its 18 tasks span six clusters and use blind pairwise judging.

Top models (higher is better)

ModelScore
Claude Fable 51154
GLM-5.21003
Grok 4.5978
GPT-5.6 Sol868
Loading Atlas data…