Atlas

Benchmarks

← All benchmarks

BAZAAR TrueSkill Mu

Agents · 2025-07-21

BAZAAR evaluates language models and baselines in a competitive double-auction market with private values and incomplete information. The public leaderboard ranks participants by multi-pass TrueSkill mu.

Top models (higher is better)

ModelScore
o37.4
Gemini 2.5 Pro7.2
Gemini 2.5 Flash7.2
Sonnet 47.1
Grok 46.9
Opus 46.9
O4 Mini6.8
DeepSeek-R1-05286.5
Qwen3-235B-A22B6.5
ChatGPT-4o Latest observed 2025-03-275.9
Llama 4 Maverick5.5
DeepSeek-V3-03245.2
Nova Pro5.2
Haiku 3.55.0
Qwen3-30B-A3B5.0
Llama 4 Scout4.9
GPT-4o Mini4.8
Mistral Small 3.2 24B Instruct 25064.8
Mistral Medium 34.7
MiniMax-Text-014.6
Phi-44.4
Kimi K2 Instruct4.4
ERNIE 4.5 300B A47B4.3
Gemma 3 27B IT4.2
Loading Atlas data…