Pointerbench
Multimodal · 2026-07-02
Pointerbench measures GUI grounding from a 1024×768 screenshot and an instruction. Its headline score is the reported public average of three equally sized 500-example subsets: Sheets, Text, and Pro. Point tasks use point-in-bounding-box accuracy; Text also includes bounding-box tasks requiring at least 90% ground-truth coverage and 70% prediction precision.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Fable 5 | 78.7 |
| Sonnet 4.6 | 69.3 |
| Opus 4.8 | 68.9 |
| GPT-5.6 Sol | 65.3 |
| GPT-5.6 Sol Pro | 65.3 |
| Opus 4.7 | 61.8 |
| GPT-5.5 | 59.9 |
| GPT-5.6 Terra Pro | 48.7 |
| GPT-5.6 Terra | 47.3 |
| GPT-5.4 | 46.2 |
| GPT-5.6 Luna | 45.3 |
| GPT-5.6 Luna Pro | 44.7 |
| GPT-5 | 23.0 |
| Kimi K2.5 | 5.7 |
| GPT-5 Mini | 4.4 |
| Kimi K2.6 | 3.7 |
| Qwen3.7-Plus | 3.1 |
| Qwen3.6-Flash-2026-04-16 | 2.7 |
| MiniMax M3 | 2.5 |
| Qwen3.6 Plus (2026-04-02) | 2.4 |
| Qwen3 VL 235B A22B Thinking | 2.1 |
| Qwen3.5-9B | 1.6 |
| Gemini 3.5 Flash | 1.5 |
| GPT-5 Nano | 1.3 |
| Kimi K2.7 Code | 1.0 |
| Grok 4.3 | 0.4 |
| Gemini 3.1 Pro Preview | 0.2 |
| Step 3.7 Flash | 0.1 |
| Grok Build 0.1 | 0.0 |
| MiMo-V2.5 | 0.0 |