OSS-Fuzz (CAISI 297 tasks)
Code · 2026-09-17
CAISI's private 297-task benchmark provides vulnerable open-source code without a bug description, crash input or patch. A 300-turn ReAct agent must find and exploit the defect to hijack the program.
Top models (higher is better)
| Model | Score |
|---|---|
| GLM-5.3 | 7.7 |