Macaron LivingBench
Agents · 2026-06-07
Macaron LivingBench evaluates personal agents on dynamic everyday scenarios involving changing users, noisy tool returns, and evolving environments. The July 2026 V1 evaluation used 40 Chinese- and English-language scenarios, up to 10 turns per case, and a score combining Need Fulfillment with Process Quality; the benchmark itself is continuously updated from production scenarios.
Top models (higher is better)
| Model | Score |
|---|---|
| Macaron-V1-Venti | 64.0 |
| Opus 4.8 | 63.8 |
| GPT-5.5 | 61.9 |
| GLM-5.2 | 60.5 |
| MiniMax M3 | 57.1 |
| Qwen3.7-Max | 56.1 |
| Gemini 3.1 Pro Preview | 52.1 |