Atlas

Benchmarks

← All benchmarks

Adversarial Robustness

Safety · 2024-07-29

Violation rate on adversarial safety prompts covering harmful, hateful, and illegal requests.

Top models (lower is better)

ModelScore
Gemini 1.5 Pro Preview 05148.0
Llama 3.1 405B Instruct10.0
Opus 313.0
Gemini 1.5 Flash Preview 051414.0
Claude 3.5 Sonnet (June 2024)16.0
GPT-4 0125 Preview20.0
Mistral Large 1.037.0
GPT-4o67.0
Loading Atlas data…