Atlas

Benchmarks

← All benchmarks

AidanBench

Chat & Writing · 2024-08-05

AidanBench stress-tests sustained open-ended generation by repeatedly asking models for novel answers to 63 open-ended prompts until outputs become incoherent or too similar to prior responses. Its headline score is the total number of valid answers across prompts, so the benchmark has no fixed score ceiling.

Top models (higher is better)

ModelScore
O16057
o3-mini4922
Gemini 2.5 Pro Preview 03-253212
GPT-4.53004
Gemini 2.0 Flash Thinking Experimental 12192735
Claude 3.5 Sonnet (Oct 2024)2673
Grok 3 Mini2285
Claude 3.7 Sonnet2170
Gemma 3 27B IT2030
o1 Preview1875
Gemini 2.0 Flash Experimental1872
Grok 31796
Gemma 3 12B IT1787
Gemini 1.5 Pro1487
GPT-4 Turbo1479
DeepSeek-R11448
o1-mini1425
Claude 3.5 Sonnet (June 2024)1361
GPT-4.11356
GPT-4o (2024-08-06)1325
GPT-4.1 Mini1318
GPT-4o (2024-11-20)1256
GPT-4 Turbo (1106 Preview)1240
GPT-4o1048
Llama 4 Maverick Instruct1036
Gemini 1.5 Flash1014
Gemma 2 27B IT1006
Mistral Large Latest (rolling alias)926
GPT-4909
DeepSeek-V3-0324892
ChatGPT-4o Latest (source-unspecified snapshot)890
Opus 3878
Haiku 3.5869
Gemma 2 9B IT843
Grok Beta823
Llama 3.3 70B Instruct776
Llama 3.1 405B Instruct759
GPT-4.1 Nano666
GPT-4o Mini640
Llama 3.1 70B Instruct578
Loading Atlas data…