Atlas

Benchmarks

← All benchmarks

BuzzBench

Chat & Writing · 2025-01-05

Humor-analysis benchmark for emotional-intelligence evaluation.

Top models (higher is better)

ModelScore
Gemini 2.5 Pro Preview 03-2571.1
ChatGPT-4o Latest observed 2025-03-2769.1
Claude 3.7 Sonnet68.2
GPT-4.167.1
DeepSeek-R162.7
Claude 3.5 Sonnet (Oct 2024)61.9
ChatGPT-4o Latest (source-unspecified snapshot)59.5
O159.3
Gemini 1.5 Pro54.9
Gemini 2.0 Flash 00154.4
Haiku 3.553.3
Grok 2 121251.8
DeepSeek-V351.6
Mistral Large 2.1 (Instruct 2411)47.5
o1-mini46.2
Llama 3.1 405B Instruct45.9
Qwen2.5 72B Instruct45.9
Llama 3.3 70B Instruct45.7
GPT-3.5 Turbo 061339.0
Phi-438.7
Llama 3.1 8B Instruct32.9
Ministral 8B Instruct 241032.3
Qwen2.5 7B Instruct31.4
Llama 3.2 3B Instruct26.6
Llama 3.2 1B Instruct9.5
Loading Atlas data…