Atlas

Benchmarks

← All benchmarks

ClockBench Accuracy

Multimodal · 2025-09-02

ClockBench Accuracy is the public leaderboard's primary time-reading metric across the benchmark's 180 analog clocks. Scores are the percentage of clock-reading answers judged correct.

Top models (higher is better)

ModelScore
GPT-5.6 Sol66.7
GPT-5.450.6
GPT-5.546.1
Qwen3-VL-235B-A22B-Instruct39.4
Claude Fable 535.0
Gemini 3.1 Pro Preview32.2
Gemini 3.5 Flash31.1
Gemini 3 Pro Preview28.9
Grok 4.521.7
Gemini 2.5 Pro18.9
GPT-5.215.0
Gemini Robotics-ER 1.515.0
Opus 4.715.0
o3-pro14.4
Qwen3 VL 235B A22B Thinking14.4
o312.2
Gemini 2.5 Flash11.1
GPT-511.1
GPT-5 Pro11.1
Mistral Medium 3.110.0
Opus 4.68.9
GPT-5 Mini8.9
Opus 4.18.3
Sonnet 4.57.2
Qwen2.5-VL-72B-Instruct6.1
Sonnet 46.1
GPT-4o5.0
GPT-5 Nano3.9
Grok 4 Fast3.9
Loading Atlas data…