Atlas

Benchmarks

← All benchmarks

ScreenSpot-Pro

Multimodal · 2024-12-19

ScreenSpot-Pro evaluates GUI grounding on high-resolution screenshots from 23 professional applications across five industries and three operating systems. Given an instruction, a model must predict the pixel location of the target UI element; a prediction counts only when it falls inside the ground-truth box.

Top models (higher is better)

ModelScore
GPT-5.286.3
Gemini 3 Pro Preview72.7
Gemini 3 Flash Preview69.1
Qwen2.5-VL-72B-Instruct53.3
Kimi-VL-A3B-Thinking-250651.0
Qwen2.5 VL 32B Instruct48.0
Sonnet 4.536.2
Qwen2.5-VL-7B-Instruct26.8
Qwen2.5-VL-3B-Instruct16.1
Gemini 2.5 Pro11.4
CogAgent7.7
Gemini 2.5 Flash3.9
GPT-5.13.5
Qwen2 VL 72B Instruct1.0
GPT-4o0.8
Loading Atlas data…