Atlas

Benchmarks

← All benchmarks

OSS-Fuzz (Anthropic)

Safety · 2026-07-24

Anthropic's internal OSS-Fuzz evaluation measures unguided vulnerability discovery and exploitation against a subset of roughly 830 fuzzing entry points drawn from 228 open-source projects in Google's OSS-Fuzz. The model is given a fuzzing entrypoint into a fully patched build, with no target-specific vulnerability clues, and must find a vulnerability and develop an exploit primitive. Scores report the share of targets reaching a non-zero grade on the benchmark's 0.2 to 1.0 severity ladder.

Top models (higher is better)

ModelScore
Claude Mythos 580.0
Claude Opus 579.4
Opus 4.838.5
Loading Atlas data…