A-Fantasia (Valid Attempts with Retries)
Games · 2026-09-30
A-Fantasia mean of chess, cube and backward-spelling error rates under the retry/valid-attempt protocol. Each task samples 100 items, allows five additional attempts for unscorable responses and divides by valid attempts only. Selects the latest run with at least 80 valid items on every task. Published rounded aggregate; variable denominators and missing attempt logs prevent an exact standard error.
Top models (lower is better)
| Model | Score |
|---|---|
| Claude Opus 5 | 30.0 |
| Opus 4.6 | 30.0 |
| Gemini 3.1 Pro Preview | 35.0 |
| Opus 4.8 | 35.0 |
| GPT-5.5 | 35.0 |
| Opus 4.1 | 35.0 |
| Opus 4.5 | 35.0 |
| GPT-5.6 Sol | 36.0 |
| Gemini 3 Flash Preview | 37.0 |
| GPT-4.5 | 38.0 |
| Opus 4 | 40.0 |
| Sonnet 4.5 | 40.0 |
| Gemini 3 Pro Preview | 42.0 |
| GPT-5.4 | 42.0 |
| GPT-5.6 Terra | 43.0 |
| Sonnet 4 | 44.0 |
| Claude 3.7 Sonnet | 45.0 |
| Claude 3.5 Sonnet (Oct 2024) | 46.0 |
| Qwen3.7-Max | 46.0 |
| Grok 3 | 47.0 |
| Sonnet 4.6 | 48.0 |
| GPT-4o | 48.0 |
| Opus 3 | 50.0 |
| GPT-5.1 | 51.0 |
| GPT-5.2 | 51.0 |
| GPT-5 Chat (2025-08-07) | 52.0 |
| Gemini 2.0 Flash 001 | 52.0 |
| GPT-4.1 | 53.0 |
| Gemini 3.1 Flash-Lite | 53.0 |
| Gemini 2.5 Flash | 55.0 |
| Kimi K2.6 | 56.0 |
| GPT-5.6 Luna | 56.0 |
| DeepSeek-V4-Pro | 57.0 |
| Kimi K2 Instruct | 57.0 |
| Haiku 4.5 | 58.0 |
| Qwen3.6 Plus (2026-04-02) | 58.0 |
| Hy4 Preview | 59.0 |
| DeepSeek V4.1 Flash | 59.0 |
| Qwen3.7-Plus | 62.0 |
| Gemini 1.5 Pro | 62.0 |