AidanBench
Chat & Writing · 2024-08-05
AidanBench stress-tests sustained open-ended generation by repeatedly asking models for novel answers to 63 open-ended prompts until outputs become incoherent or too similar to prior responses. Its headline score is the total number of valid answers across prompts, so the benchmark has no fixed score ceiling.