ClawGym-Bench
Agents · 2026-05-15
ClawGym-Bench evaluates Claw-style personal agents on 200 long-horizon tasks in realistic local, stateful workspaces. Its 156 deterministic tasks use code verification, while 44 hybrid tasks combine code checks at 70% weight with rubric-based judgment at 30%.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.5 | 82.5 |
| Opus 4.8 | 80.5 |
| Macaron-V1-Venti | 77.7 |
| Gemini 3.1 Pro Preview | 77.5 |
| MiniMax M3 | 76.2 |
| Qwen3.7-Max | 75.7 |
| GLM-5.2 | 74.6 |