Atlas

Benchmarks

← All benchmarks

CadQueryEval (Corrected Geometry)

Code · 2026-09-30

CadQueryEval's final September 2026 25-task release. One-shot CadQuery generation; all execution and mesh geometry checks must pass. The corrected scorer accepts either documented interpretation of task 6 and uses trimesh watertightness. Kept separate from the earlier reference/scorer. Scores come from the final README tables; the repository's older results.json is a different release.

Top models (higher is better)

ModelScore
Claude Opus 5.5100.0
GPT-5.6 Sol Pro100.0
GPT-6 Astra100.0
GPT-6 Sol100.0
Gemini 3.8 Flash96.0
GPT-6 Luna96.0
GLM-5.3 Flash92.0
Gemini 3.1 Pro Preview88.0
GPT-5.6 Luna88.0
GPT-5.6 Luna Pro88.0
GPT-5.6 Sol88.0
Grok 4.788.0
Gemini 3.7 Flash84.0
GPT-5.6 Terra Pro84.0
Kimi K384.0
Muse Spark 1.384.0
Claude Fable 580.0
Claude Opus 580.0
Qwen3.7-Max80.0
Claude Fable 5.176.0
DeepSeek V4.1 Flash76.0
GLM-5.376.0
GPT-5.576.0
MiMo-V2.6-Pro76.0
Gemini 3.5 Flash72.0
Gemini 3.6 Flash72.0
Grok 4.572.0
Grok 4.672.0
Opus 4.668.0
Opus 4.868.0
Sonnet 568.0
GPT-5.6 Terra68.0
Grok 4.368.0
MiMo-V2.6-Flash68.0
Muse Spark 1.164.0
Opus 4.760.0
GLM-5.160.0
Kimi K2.660.0
Muse Spark 1.260.0
Gemini 3 Pro Preview56.0
Loading Atlas data…