Atlas

Benchmarks

← All benchmarks

CadQueryEval — Code Execution

Code · 2026-01-10

Percentage of CadQueryEval tasks for which the generated CadQuery Python code executes successfully within the Docker sandbox's 60-second limit.

Top models (higher is better)

ModelScore
Gemini 3.1 Pro Preview100.0
GPT-5.6 Sol Pro100.0
GPT-5.596.0
GPT-5.6 Sol96.0
GPT-5.6 Luna92.0
GPT-5.6 Luna Pro92.0
GPT-5.6 Terra Pro92.0
Opus 4.888.0
DeepSeek-V4-Pro88.0
GPT-5.6 Terra88.0
Grok 4.588.0
Kimi K388.0
Qwen3.7-Max88.0
Claude Fable 584.0
Opus 4.584.0
Opus 4.684.0
Sonnet 584.0
Gemini 3.5 Flash80.0
GPT-5 Mini80.0
Grok 4.380.0
Kimi K2.680.0
MiniMax M380.0
Muse Spark 1.180.0
Opus 4.776.0
Sonnet 4.676.0
Gemini 3 Pro Preview76.0
GLM-5.176.0
GLM-5.276.0
Kimi K2.576.0
Qwen3.7-Plus76.0
Gemini 3 Flash Preview72.0
GPT-572.0
GPT-5.472.0
Opus 468.0
Sonnet 468.0
GPT-5.168.0
GPT-5.268.0
O4 Mini68.0
Opus 4.164.0
Hy3 preview64.0
Loading Atlas data…