Sonnet 4.6
Anthropic · 2026-02-17
Sonnet 4.6 is Anthropic's balanced mid-tier Claude model, released as a hybrid-reasoning model supporting both extended and adaptive thinking. It improved broadly over Sonnet 4.5 across coding, computer use, long-context reasoning, agent planning, and knowledge work, with a 1M-token context window.
Benchmark scores
| Benchmark | Score |
|---|---|
| A-Fantasia | 52.3 |
| A-Fantasia Backwards Spelling | 39.0 |
| A-Fantasia Chess | 26.0 |
| A-Fantasia Cube Rotation | 92.0 |
| AA-Briefcase Elo | 1076 |
| AA-LCR | 70.7 |
| AA-Omniscience Index | 12.4 |
| Agent Red Teaming k=100 (CAIS) | 54.8 |
| Agents' Last Exam ALE-CLI Pass Rate | 13.3 |
| Agents' Last Exam ALE-CLI Score | 32.0 |
| AHK-Eval | 97.2 |
| AHK-Eval Algorithms | 66.7 |
| AHK-Eval Data Structures | 50.0 |
| AHK-Eval Date and Time | 100.0 |
| AHK-Eval Easy | 75.0 |
| AHK-Eval Hard | 58.3 |
| AHK-Eval Hidden Cases | 99.4 |
| AHK-Eval Mid | 75.0 |
| AHK-Eval Numbers | 66.7 |
| AHK-Eval Parse Success | 97.2 |
| AHK-Eval Regex | 66.7 |
| AHK-Eval Strings | 66.7 |
| AIME 2024 (avg@8) | 93.3 |
| AIME 2024-2025 | 92.3 |
| AIME 2025 | 95.6 |
| AIME 2025 (avg@8) | 91.3 |
| ALE-Bench | 1768 |
| Antidote: Everyday Edition | 1013 |
| APEX (Mercor) | 10.5 |
| APEX-Agents | 23.7 |