Atlas

Benchmarks

← All benchmarks

Macaron ChatBench

Chat & Writing · 2026-07-21

Macaron ChatBench evaluates 46 de-identified cases from real multi-turn conversations, split evenly between interaction quality and task understanding and completion. Each model-case pair is judged three times by a privately deployed GLM-5.2 using case-specific 1-to-5 rubrics, then averaged and normalized to a 100-point scale.

Top models (higher is better)

ModelScore
Macaron-V1-Venti58.3
GPT-5.555.5
GLM-5.254.5
Opus 4.852.8
Qwen3.7-Max52.5
Gemini 3.1 Pro Preview52.0
MiniMax M349.1
Loading Atlas data…