Atlas

Benchmarks

← All benchmarks

Pointerbench — Text

Multimodal · 2026-07-02

Pointerbench-Text is a 500-example synthetic GUI-grounding benchmark for words, characters, punctuation, caret positions, interface text, text bounding boxes, and invoice fields. Point tasks use point-in-bounding-box accuracy; bounding-box tasks require at least 90% target coverage and 70% prediction precision.

Top models (higher is better)

ModelScore
Claude Fable 558.4
Sonnet 4.647.8
Opus 4.844.4
Opus 4.736.2
GPT-5.6 Sol Pro32.0
GPT-5.531.4
GPT-5.6 Sol30.0
GPT-5.425.0
GPT-5.6 Terra22.0
GPT-5.6 Luna20.0
GPT-5.6 Terra Pro20.0
GPT-5.6 Luna Pro18.0
GPT-513.6
Kimi K2.53.4
Kimi K2.62.8
Kimi K2.7 Code1.8
GPT-5 Mini1.4
Qwen3.6-Flash-2026-04-161.2
Gemini 3.5 Flash0.6
Qwen3.6 Plus (2026-04-02)0.6
Qwen3.7-Plus0.6
Gemini 3.1 Pro Preview0.4
MiniMax M30.4
Qwen3 VL 235B A22B Thinking0.4
GPT-5 Nano0.2
Qwen3.5-9B0.2
Grok 4.30.0
Grok Build 0.10.0
MiMo-V2.50.0
Step 3.7 Flash0.0
Loading Atlas data…