ScreenSpot-Pro
Multimodal · 2024-12-19
ScreenSpot-Pro evaluates GUI grounding on high-resolution screenshots from 23 professional applications across five industries and three operating systems. Given an instruction, a model must predict the pixel location of the target UI element; a prediction counts only when it falls inside the ground-truth box.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.2 | 86.3 |
| Gemini 3 Pro Preview | 72.7 |
| Gemini 3 Flash Preview | 69.1 |
| Qwen2.5-VL-72B-Instruct | 53.3 |
| Kimi-VL-A3B-Thinking-2506 | 51.0 |
| Qwen2.5 VL 32B Instruct | 48.0 |
| Sonnet 4.5 | 36.2 |
| Qwen2.5-VL-7B-Instruct | 26.8 |
| Qwen2.5-VL-3B-Instruct | 16.1 |
| Gemini 2.5 Pro | 11.4 |
| CogAgent | 7.7 |
| Gemini 2.5 Flash | 3.9 |
| GPT-5.1 | 3.5 |
| Qwen2 VL 72B Instruct | 1.0 |
| GPT-4o | 0.8 |