GPQA Diamond (Artificial Analysis avg@5)
General QA · 2023-11-20
Artificial Analysis' zero-shot GPQA Diamond protocol averages pass@1 over five independent attempts for each of 198 questions (990 scored attempts). This repeated-run aggregate is separated from single-attempt GPQA observations so it does not receive a binomial 198-item noise model.
Top models (higher is better)
| Model | Score |
|---|---|
| Gemini 3.1 Pro Preview | 94.1 |
| GPT-5.6 Sol | 94.1 |
| Claude Opus 5 | 93.7 |
| GPT-5.5 | 93.5 |
| Kimi K3 | 93.5 |
| Grok 4.5 | 93.1 |
| MiniMax M3 | 92.9 |
| Gemini 3.6 Flash | 92.8 |
| Claude Fable 5 | 92.6 |
| GPT-5.6 Terra | 92.5 |
| Qwen3.7-Max | 92.3 |
| Gemini 3.5 Flash | 92.2 |
| Opus 4.8 | 92.0 |
| GPT-5.4 | 92.0 |
| GPT-5.3-Codex | 91.5 |
| Opus 4.7 | 91.4 |
| Sonnet 5 | 91.1 |
| GPT-5.6 Luna | 91.1 |
| Grok 4.20 | 91.1 |
| Kimi K2.6 | 91.1 |
| DeepSeek-V4-Flash-0731 | 90.8 |
| Gemini 3 Pro Preview | 90.8 |
| DeepSeek-V4-Pro | 90.5 |
| GPT-5.2 | 90.3 |
| Grok 4.3 | 90.1 |
| Qwen3.7-Plus | 90.0 |
| GPT-5.2-Codex | 89.9 |
| Gemini 3 Flash Preview | 89.8 |
| Muse Spark 1.1 | 89.8 |
| Hy3 | 89.7 |
| Opus 4.6 | 89.6 |
| Kimi K2.7 Code | 89.6 |
| GLM-5.2 | 89.5 |
| Grok Build 0.1 | 89.5 |
| Inkling-Small | 89.5 |
| DeepSeek-V4-Flash | 89.4 |
| Qwen3.5 397B A17B | 89.3 |
| Nex-N2-Pro | 89.2 |
| Qwen3.6-Max-Preview | 88.8 |
| Muse Spark | 88.4 |