Atlas

Benchmarks

← All benchmarks

A-Fantasia (Valid Attempts with Retries)

Games · 2026-09-30

A-Fantasia mean of chess, cube and backward-spelling error rates under the retry/valid-attempt protocol. Each task samples 100 items, allows five additional attempts for unscorable responses and divides by valid attempts only. Selects the latest run with at least 80 valid items on every task. Published rounded aggregate; variable denominators and missing attempt logs prevent an exact standard error.

Top models (lower is better)

ModelScore
Claude Opus 530.0
Opus 4.630.0
Gemini 3.1 Pro Preview35.0
Opus 4.835.0
GPT-5.535.0
Opus 4.135.0
Opus 4.535.0
GPT-5.6 Sol36.0
Gemini 3 Flash Preview37.0
GPT-4.538.0
Opus 440.0
Sonnet 4.540.0
Gemini 3 Pro Preview42.0
GPT-5.442.0
GPT-5.6 Terra43.0
Sonnet 444.0
Claude 3.7 Sonnet45.0
Claude 3.5 Sonnet (Oct 2024)46.0
Qwen3.7-Max46.0
Grok 347.0
Sonnet 4.648.0
GPT-4o48.0
Opus 350.0
GPT-5.151.0
GPT-5.251.0
GPT-5 Chat (2025-08-07)52.0
Gemini 2.0 Flash 00152.0
GPT-4.153.0
Gemini 3.1 Flash-Lite53.0
Gemini 2.5 Flash55.0
Kimi K2.656.0
GPT-5.6 Luna56.0
DeepSeek-V4-Pro57.0
Kimi K2 Instruct57.0
Haiku 4.558.0
Qwen3.6 Plus (2026-04-02)58.0
Hy4 Preview59.0
DeepSeek V4.1 Flash59.0
Qwen3.7-Plus62.0
Gemini 1.5 Pro62.0
Loading Atlas data…