Atlas

Benchmarks

← All benchmarks

ActiveVision with tools

Multimodal · 2026-07-17

ActiveVision with tools is the paper's optional coding-agent ablation on the same 85 image questions. Codex or Claude Code agents operate in fresh sandboxes containing only the image and question, where they may write and run vision code. Scores report exact-match accuracy; this diagnostic is kept separate from the no-tool headline protocol.

Top models (higher is better)

ModelScore
Claude Fable 550.6
GPT-5.537.6
Opus 4.824.7
Loading Atlas data…