Market-Bench
Professional Work · 2025-12-13
Market-Bench evaluates whether models can generate executable quantitative-trading backtesters for three progressively harder strategies. Generated CSV metrics are compared with a verifiable reference backtester using mean absolute error; the benchmark models consumed liquidity through persistent synthetic order books and introduces exchange delay in the two harder strategies.
Top models (lower is better)
| Model | Score |
|---|---|
| Grok 4 | 443 |
| GPT-5.2 | 969 |
| Gemini 3 Pro Preview | 1744 |
| GPT-5.1-Codex-Max | 4243 |
| DeepSeek-V3.2 | 4576 |
| Sonnet 4.5 | 5127 |
| Opus 4.5 | 6040 |
| Command A | 6562 |
| Nova Premier | 7740 |
| Llama 3.1 Nemotron Ultra 253B V1 | 9674 |
| Llama 4 Maverick | 10202 |
| Mistral Large 3 675B Instruct 2512 | 30606 |
| Qwen3-Max (2025-09-23) | 159144490 |