Atlas

Benchmarks

← All benchmarks

OSS-Fuzz (CAISI 297 tasks)

Code · 2026-09-17

CAISI's private 297-task benchmark provides vulnerable open-source code without a bug description, crash input or patch. A 300-turn ReAct agent must find and exploit the defect to hijack the program.

Top models (higher is better)

ModelScore
GLM-5.37.7
Loading Atlas data…