Macaron ChatBench
Chat & Writing · 2026-07-21
Macaron ChatBench evaluates 46 de-identified cases from real multi-turn conversations, split evenly between interaction quality and task understanding and completion. Each model-case pair is judged three times by a privately deployed GLM-5.2 using case-specific 1-to-5 rubrics, then averaged and normalized to a 100-point scale.
Top models (higher is better)
| Model | Score |
|---|---|
| Macaron-V1-Venti | 58.3 |
| GPT-5.5 | 55.5 |
| GLM-5.2 | 54.5 |
| Opus 4.8 | 52.8 |
| Qwen3.7-Max | 52.5 |
| Gemini 3.1 Pro Preview | 52.0 |
| MiniMax M3 | 49.1 |