Atlas

Benchmarks

← All benchmarks

Macaron LivingBench

Agents · 2026-06-07

Macaron LivingBench evaluates personal agents on dynamic everyday scenarios involving changing users, noisy tool returns, and evolving environments. The July 2026 V1 evaluation used 40 Chinese- and English-language scenarios, up to 10 turns per case, and a score combining Need Fulfillment with Process Quality; the benchmark itself is continuously updated from production scenarios.

Top models (higher is better)

ModelScore
Macaron-V1-Venti64.0
Opus 4.863.8
GPT-5.561.9
GLM-5.260.5
MiniMax M357.1
Qwen3.7-Max56.1
Gemini 3.1 Pro Preview52.1
Loading Atlas data…