Atlas

Benchmarks

← All benchmarks

Pointerbench

Multimodal · 2026-07-02

Pointerbench measures GUI grounding from a 1024×768 screenshot and an instruction. Its headline score is the reported public average of three equally sized 500-example subsets: Sheets, Text, and Pro. Point tasks use point-in-bounding-box accuracy; Text also includes bounding-box tasks requiring at least 90% ground-truth coverage and 70% prediction precision.

Top models (higher is better)

ModelScore
Claude Fable 578.7
Sonnet 4.669.3
Opus 4.868.9
GPT-5.6 Sol65.3
GPT-5.6 Sol Pro65.3
Opus 4.761.8
GPT-5.559.9
GPT-5.6 Terra Pro48.7
GPT-5.6 Terra47.3
GPT-5.446.2
GPT-5.6 Luna45.3
GPT-5.6 Luna Pro44.7
GPT-523.0
Kimi K2.55.7
GPT-5 Mini4.4
Kimi K2.63.7
Qwen3.7-Plus3.1
Qwen3.6-Flash-2026-04-162.7
MiniMax M32.5
Qwen3.6 Plus (2026-04-02)2.4
Qwen3 VL 235B A22B Thinking2.1
Qwen3.5-9B1.6
Gemini 3.5 Flash1.5
GPT-5 Nano1.3
Kimi K2.7 Code1.0
Grok 4.30.4
Gemini 3.1 Pro Preview0.2
Step 3.7 Flash0.1
Grok Build 0.10.0
MiMo-V2.50.0
Loading Atlas data…