Big Sleep Vulnerability Evaluation
Safety · 2026-07-21
Google’s private Big Sleep evaluation measures pass@1 vulnerability-discovery performance in the Big Sleep agent without safety guardrails. The release does not disclose the task set, its size, or enough methodology for cross-source comparison, so Atlas retains it as a diagnostic.
Top models (higher is better)
| Model | Score |
|---|---|
| Gemini 3.5 Flash Cyber | 72.0 |
| Gemini 3.6 Flash | 42.0 |
| Gemini 3.5 Flash | 36.0 |