Adversarial Robustness
Safety · 2024-07-29
Violation rate on adversarial safety prompts covering harmful, hateful, and illegal requests.
Top models (lower is better)
| Model | Score |
|---|---|
| Gemini 1.5 Pro Preview 0514 | 8.0 |
| Llama 3.1 405B Instruct | 10.0 |
| Opus 3 | 13.0 |
| Gemini 1.5 Flash Preview 0514 | 14.0 |
| Claude 3.5 Sonnet (June 2024) | 16.0 |
| GPT-4 0125 Preview | 20.0 |
| Mistral Large 1.0 | 37.0 |
| GPT-4o | 67.0 |