Atlas

Benchmarks

← All benchmarks

CorpFin v2 - Max Fitting Context

Professional Work · 2025-01-27

The Max Fitting Context split of CorpFin v2. This child benchmark separates a source-reported subtask or subtrack from the parent aggregate so scores at different grains do not share one benchmark_slug.

Top models (higher is better)

ModelScore
Claude Fable 577.0
Claude Opus 576.0
Muse Spark 1.174.4
Kimi K372.7
Qwen3 Max Thinking (2026-01-23)70.6
Inkling-Small70.4
GPT-5.570.0
Opus 4.869.8
MiniMax M369.7
Sonnet 569.1
Grok 4.368.8
GLM-5.268.4
Kimi K2.568.4
Qwen3.6-Max-Preview68.2
Opus 4.767.8
Inkling67.6
Grok 4.567.5
Kimi K2.667.2
Qwen3.5 Plus (2026-02-15)67.2
Opus 4.667.1
Nemotron 3 Ultra 550B A55B67.1
Gemini 3 Pro Preview67.0
Gemini 3.1 Pro Preview66.9
Gemini 3.5 Flash66.7
GLM-5.166.3
GPT-5.6 Terra66.3
Sonnet 4.666.1
Qwen3.5-Flash65.7
Qwen3.7-Max65.7
Opus 4.565.6
Gemini 3.6 Flash65.4
GPT-5.6 Sol65.2
GPT-5.464.9
Qwen3.6 27B64.9
Grok 4.1 Fast64.8
GPT-5.6 Luna64.1
GPT-5.263.9
Mistral Large 3 675B Instruct 251263.8
Gemini 3 Flash Preview63.5
GPT-5.163.4
Loading Atlas data…