Atlas

Benchmarks

← All benchmarks

CadQueryEval — Bounding Box

Code · 2026-01-10

Percentage of CadQueryEval tasks whose generated geometry matches the reference bounding-box dimensions within 1.0 mm.

Top models (higher is better)

ModelScore
Gemini 3.1 Pro Preview100.0
GPT-5.6 Sol Pro100.0
GPT-5.6 Luna Pro92.0
GPT-5.6 Sol92.0
GPT-5.6 Luna88.0
GPT-5.6 Terra Pro88.0
Grok 4.588.0
Kimi K388.0
Qwen3.7-Max88.0
Claude Fable 584.0
Opus 4.680.0
Sonnet 580.0
GPT-5.580.0
Grok 4.380.0
Opus 4.876.0
Gemini 3.5 Flash76.0
GPT-5.6 Terra76.0
MiniMax M376.0
DeepSeek-V4-Pro72.0
Kimi K2.572.0
Muse Spark 1.172.0
Opus 4.768.0
GLM-5.168.0
GPT-5 Mini68.0
Kimi K2.668.0
Opus 4.564.0
Gemini 3 Pro Preview64.0
Gemini 3 Flash Preview60.0
GLM-5.260.0
Qwen3.7-Plus60.0
Gemini 3.1 Flash-Lite56.0
GPT-5.456.0
Sonnet 4.552.0
Sonnet 4.652.0
DeepSeek-V4-Flash52.0
Hy3 preview52.0
O4 Mini52.0
Claude 3.5 Sonnet (Oct 2024)48.0
Opus 4.148.0
GPT-548.0
Loading Atlas data…