Atlas

Benchmarks

← All benchmarks

Roboflow Vision Evals (September 2026 protocol)

Multimodal

Unweighted mean of six low-tier vision task headlines: detection mAP@50, OCR similarity, and Gemini 3.5 Flash temperature-zero judged accuracy for counting, identification, extraction and reasoning. Newer configurations average three runs; older configurations retain a single run. The September marker distinguishes this observed protocol from the earlier strict-match snapshot, not an upstream version number. The current pages do not state atomic sample counts.

Top models (higher is better)

ModelScore
GPT-6 Astra86.6
Gemini 3.5 Flash86.0
Claude Opus 5.585.5
Gemini 3.7 Flash85.2
Gemini 3.8 Flash85.1
Gemini 3.1 Pro Preview83.3
Gemini 3.6 Flash83.0
Claude Fable 5.181.3
GPT-6 Sol80.7
Muse Spark 1.180.5
Muse Spark 1.280.5
Muse Spark 1.379.8
GPT-5.6 Sol79.0
Claude Fable 578.7
Claude Opus 578.3
Gemini 3 Flash Preview74.9
GPT-5.574.8
GPT-5.6 Luna73.8
GPT-5.6 Terra73.8
Grok 4.771.9
Muse Glimmer 30B70.8
Gemini 3.5 Flash-Lite70.3
Qwen3.8 Flash68.8
Grok 4.668.7
Opus 4.868.7
GPT-6 Luna68.6
Qwen3.7-Plus67.4
Kimi K366.5
Sonnet 566.4
GLM-5.3 Flash66.3
Gemini 2.5 Pro66.0
Qwen3-VL-235B-A22B-Instruct65.9
Grok 4.565.8
GLM 5V Turbo65.3
GPT-5.4 Mini64.7
Qwen3.5-27B64.4
Kimi K2.662.9
MiMo-V2.6-Pro62.5
Qwen3.7 Flash61.5
MiMo-V2.6-Flash60.6
Loading Atlas data…