Atlas

Benchmarks

← All benchmarks

PolicyBench US Exact Match

Professional Work · 2026-06-10

PolicyBench tests whether language models can calculate household taxes and benefits without tools. Its US exact-match leaderboard covers 100 households and 18 tax, credit, benefit, and eligibility outputs; amount predictions must match within one dollar and binary predictions exactly, with output groups weighted by estimated household impact.

Top models (higher is better)

ModelScore
GPT-5.6 Sol88.7
GPT-5.6 Luna84.5
GPT-5.583.5
GPT-5.6 Terra83.4
Grok 4.580.9
Claude Fable 579.9
Gemini 3.1 Pro Preview77.9
Opus 4.777.4
Grok 4.377.2
Sonnet 4.677.1
Gemini 3 Flash Preview76.9
Gemini 3.5 Flash76.2
Grok Build 0.176.1
Gemini 3.1 Flash-Lite Preview76.1
DeepSeek-V4-Pro76.1
Qwen3.7-Max73.6
GLM-5.273.1
Opus 4.872.6
MiniMax M372.4
Haiku 4.571.7
GPT-5.4 Mini70.5
Sonnet 569.4
Kimi K2.664.6
GPT-5.4 Nano62.3
Loading Atlas data…