Atlas

Benchmarks

← All benchmarks

ClawGym-Bench

Agents · 2026-05-15

ClawGym-Bench evaluates Claw-style personal agents on 200 long-horizon tasks in realistic local, stateful workspaces. Its 156 deterministic tasks use code verification, while 44 hybrid tasks combine code checks at 70% weight with rubric-based judgment at 30%.

Top models (higher is better)

ModelScore
GPT-5.582.5
Opus 4.880.5
Macaron-V1-Venti77.7
Gemini 3.1 Pro Preview77.5
MiniMax M376.2
Qwen3.7-Max75.7
GLM-5.274.6
Loading Atlas data…