ActiveVision
Multimodal · 2026-07-17
ActiveVision evaluates iterative visual reasoning on 85 photorealistic image questions across 17 tasks in distributed scanning, sequential traversal, and visual attribute transfer. Under the benchmark's primary protocol, models receive one image and question per request and use pure chain-of-thought with no external tools. Scores report exact-match accuracy after separator-insensitive answer normalization.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.5 | 10.6 |
| Gemini 3.5 Flash | 8.2 |
| Claude Fable 5 | 5.9 |
| Gemini 3.1 Pro Preview | 5.9 |
| Opus 4.7 | 4.7 |
| Opus 4.8 | 2.4 |