Atlas

Benchmarks

← All benchmarks

Big Sleep Vulnerability Evaluation

Safety · 2026-07-21

Google’s private Big Sleep evaluation measures pass@1 vulnerability-discovery performance in the Big Sleep agent without safety guardrails. The release does not disclose the task set, its size, or enough methodology for cross-source comparison, so Atlas retains it as a diagnostic.

Top models (higher is better)

ModelScore
Gemini 3.5 Flash Cyber72.0
Gemini 3.6 Flash42.0
Gemini 3.5 Flash36.0
Loading Atlas data…