Atlas

Benchmarks

← All benchmarks

HijackEval Public Attacks (best of 31)

Safety · 2026-07-17

This HijackEval protocol reports held-out attack success using the best public attack selected for each model, attacker task, and vector from 31 candidates on a separate development set. Results cover four malicious tasks and three vectors and count both partial and full malicious-task completion.

Top models (lower is better)

ModelScore
GLM-5.20.0
Opus 4.70.1
GPT-5.50.5
Kimi K2.61.9
DeepSeek-V4-Pro15.1
Loading Atlas data…