Atlas

Benchmarks

← All benchmarks

BALROG

Games · 2024-11-20

BALROG evaluates agents on long-horizon games with widely varying difficulty. Its environments test planning, spatial reasoning, exploration, and learning from interaction.

Top models (higher is better)

ModelScore
Gemini 3 Pro Preview58.1
Gemini 3 Flash Preview48.1
Grok 443.6
Opus 4.543.5
Gemini 2.5 Pro Experimental 03-2543.3
DeepSeek-R134.9
Gemini 2.5 Flash33.5
GPT-532.8
Claude 3.5 Sonnet (Oct 2024)32.6
GPT-4o32.3
Haiku 4.531.2
Grok 329.5
Reka Flash 329.2
Llama 3.1 70B Instruct27.9
Llama 3.2 90B Vision Instruct27.3
Llama 3.3 70B Instruct23.0
Gemini 1.5 Pro 00221.0
DeepSeek-R1-Distill-Qwen-32B19.5
Haiku 3.519.3
Mistral NeMo Instruct 240717.6
GPT-4o Mini17.4
Llama 3.2 11B Vision Instruct16.8
Qwen2.5 72B Instruct16.2
Llama 3.1 8B Instruct15.1
Gemini 1.5 Flash 00214.6
Qwen2 VL 72B Instruct12.8
Phi-411.6
Llama 3.2 3B Instruct10.1
Qwen2.5 7B Instruct7.8
Llama 3.2 1B Instruct6.6
Qwen2 VL 7B Instruct3.7
Loading Atlas data…