Cybench
Agents · 2024-08-15
A cybersecurity agent benchmark measuring autonomous vulnerability discovery and exploitation across sandboxed challenges.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Mythos Preview | 100.0 |
| GPT-5 | 73.5 |
| gpt-oss-120b | 49.5 |
| Opus 4 | 46.9 |
| DeepSeek-V3.1 | 40.0 |
| DeepSeek-R1-0528 | 35.5 |
| DeepSeek-R1 | 16.7 |