PolicyBench US Exact Match
Professional Work · 2026-06-10
PolicyBench tests whether language models can calculate household taxes and benefits without tools. Its US exact-match leaderboard covers 100 households and 18 tax, credit, benefit, and eligibility outputs; amount predictions must match within one dollar and binary predictions exactly, with output groups weighted by estimated household impact.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.6 Sol | 88.7 |
| GPT-5.6 Luna | 84.5 |
| GPT-5.5 | 83.5 |
| GPT-5.6 Terra | 83.4 |
| Grok 4.5 | 80.9 |
| Claude Fable 5 | 79.9 |
| Gemini 3.1 Pro Preview | 77.9 |
| Opus 4.7 | 77.4 |
| Grok 4.3 | 77.2 |
| Sonnet 4.6 | 77.1 |
| Gemini 3 Flash Preview | 76.9 |
| Gemini 3.5 Flash | 76.2 |
| Grok Build 0.1 | 76.1 |
| Gemini 3.1 Flash-Lite Preview | 76.1 |
| DeepSeek-V4-Pro | 76.1 |
| Qwen3.7-Max | 73.6 |
| GLM-5.2 | 73.1 |
| Opus 4.8 | 72.6 |
| MiniMax M3 | 72.4 |
| Haiku 4.5 | 71.7 |
| GPT-5.4 Mini | 70.5 |
| Sonnet 5 | 69.4 |
| Kimi K2.6 | 64.6 |
| GPT-5.4 Nano | 62.3 |