Atlas

Benchmarks

← All benchmarks

CadQueryEval

Code · 2026-01-10

CadQueryEval measures the percentage of 25 natural-language CAD tasks for which a model's one-shot CadQuery Python program executes, exports an STL, and passes every geometry check against a reference model: watertightness, component count, bounding-box dimensions, volume, Chamfer distance, and 95th-percentile Hausdorff distance.

Top models (higher is better)

ModelScore
GPT-5.6 Sol Pro84.0
Gemini 3.1 Pro Preview80.0
GPT-5.6 Luna76.0
GPT-5.6 Luna Pro76.0
GPT-5.6 Sol76.0
Kimi K376.0
GPT-5.6 Terra Pro72.0
Qwen3.7-Max72.0
Claude Fable 564.0
GPT-5.564.0
Opus 4.660.0
Sonnet 560.0
Gemini 3.5 Flash60.0
Grok 4.360.0
Grok 4.560.0
Opus 4.856.0
GLM-5.156.0
GPT-5.6 Terra56.0
Kimi K2.656.0
Muse Spark 1.156.0
Opus 4.752.0
GPT-5 Mini52.0
Hy3 preview52.0
Qwen3.7-Plus52.0
Gemini 3 Pro Preview48.0
Kimi K2.548.0
Sonnet 4.544.0
GLM-5.244.0
MiniMax M344.0
O144.0
Opus 4.540.0
Gemini 3 Flash Preview40.0
DeepSeek-V4-Flash36.0
Gemini 3.1 Flash-Lite36.0
GPT-536.0
GPT-5.136.0
GPT-5.236.0
GPT-5.436.0
o336.0
Sonnet 4.632.0
Loading Atlas data…