Atlas

Models

← All models

Grok 2

xAI · 2024-08-13

Grok 2 is xAI's second-generation flagship model, released as a hosted beta in August 2024 with substantially improved chat, coding, reasoning, tool use, and visual understanding over Grok 1.5. It launched alongside Grok 2 Mini and later reached xAI's enterprise API; the separately released August 2025 open-weight checkpoint is a later artifact and is not represented by this row.

Benchmark scores

BenchmarkScore
AIME 2024 (avg@8)19.6
AIME 2024-202515.2
AIME 2025 (avg@8)10.8
Capability128.2
Confabulations Leaderboard Confabulation Rate25.7
Confabulations Leaderboard Non-Response Rate14.5
Confabulations Leaderboard Weighted Score20.1
CorpFin v251.1
CorpFin v2 - Exact Pages57.2
CorpFin v2 - Max Fitting Context45.2
CorpFin v2 - Shared Max Context50.9
Elimination Game TrueSkill Mu3.4
GPQA Diamond52.0
GPQA Diamond (Vals protocol average)50.8
LiveCodeBench (Vals AI public dataset)38.7
LiveCodeBench (Vals Public) - Easy85.7
LiveCodeBench (Vals Public) - Hard3.4
LiveCodeBench (Vals Public) - Medium26.9
LLM Deception Disinformation Effectiveness1.0
LLM Deception Vulnerability Score0.8
LLM Divergent Thinking Repeat Rate5.6
LLM Divergent Thinking Score4.5
MATH-50078.4
MathVista (mini)69.0
MedQA Demographic Bias (Vals)92.3
MedQA Demographic Bias (Vals) - Asian92.0
MedQA Demographic Bias (Vals) - Black92.2
MedQA Demographic Bias (Vals) - Hispanic91.7
MedQA Demographic Bias (Vals) - Indigenous91.6
MedQA Demographic Bias (Vals) - Unbiased94.0
Loading Atlas data…