Atlas

Benchmarks

← All benchmarks

Market-Bench

Professional Work · 2025-12-13

Market-Bench evaluates whether models can generate executable quantitative-trading backtesters for three progressively harder strategies. Generated CSV metrics are compared with a verifiable reference backtester using mean absolute error; the benchmark models consumed liquidity through persistent synthetic order books and introduces exchange delay in the two harder strategies.

Top models (lower is better)

ModelScore
Grok 4443
GPT-5.2969
Gemini 3 Pro Preview1744
GPT-5.1-Codex-Max4243
DeepSeek-V3.24576
Sonnet 4.55127
Opus 4.56040
Command A6562
Nova Premier7740
Llama 3.1 Nemotron Ultra 253B V19674
Llama 4 Maverick10202
Mistral Large 3 675B Instruct 251230606
Qwen3-Max (2025-09-23)159144490
Loading Atlas data…