Atlas

Benchmarks

← All benchmarks

ExploitGym Userspace (CAISI 498 scored tasks)

Code · 2026-09-17

CAISI's GLM-5.3 userspace exploit-development evaluation with a 200-turn ReAct agent. The published denominator is 498, rather than the full suite's 502 tasks; the omitted-task identities are unspecified, so this cohort remains separate.

Top models (higher is better)

ModelScore
GLM-5.39.4
Loading Atlas data…