Atlas

Benchmarks

← All benchmarks

SciCode

Science · 2024-07-12

SciCode is a scientist-curated benchmark with 80 research coding problems decomposed into 338 subproblems; the standard test split used by the reported leaderboard has 288 subproblems from 65 main problems. It measures whether models can generate executable scientific code from research-style specifications.

Top models (higher is better)

ModelScore
Claude Fable 560.2
Gemini 3.1 Pro Preview58.9
Kimi K358.7
Muse Spark 1.158.2
GPT-5.6 Sol56.9
GPT-5.456.6
GPT-5.556.1
Claude Opus 555.7
Opus 4.754.5
Grok 4.554.1
GPT-5.6 Terra53.9
Sonnet 553.6
Opus 4.853.5
Kimi K2.653.5
Gemini 3.5 Flash53.1
Gemini 3.6 Flash52.7
GPT-5.6 Luna52.5
Muse Spark51.5
GLM-5.250.5
Grok Build 0.150.2
MiMo-V2.5-Pro50.2
DeepSeek-V4-Pro50.0
GPT-5.4 Mini49.9
Qwen3.7-Max48.8
GPT-5.5 Instant (2026-06-25 hosted snapshot)48.6
Kimi K2.7 Code47.5
Grok 4.347.3
MiniMax M2.747.0
GPT-5.4 Nano46.9
Sonnet 4.646.8
Qwen3.7-Plus45.5
MiniMax M345.4
DeepSeek-V4-Flash44.9
GLM-5.143.8
Gemma 4 31B IT43.4
Haiku 4.543.3
GPT-5.143.3
MiMo-V2.543.1
Gemini 2.5 Pro42.8
Nova 2 Pro (Preview)42.7
Loading Atlas data…