HijackEval Public Attacks (best of 31)
Safety · 2026-07-17
This HijackEval protocol reports held-out attack success using the best public attack selected for each model, attacker task, and vector from 31 candidates on a separate development set. Results cover four malicious tasks and three vectors and count both partial and full malicious-task completion.
Top models (lower is better)
| Model | Score |
|---|---|
| GLM-5.2 | 0.0 |
| Opus 4.7 | 0.1 |
| GPT-5.5 | 0.5 |
| Kimi K2.6 | 1.9 |
| DeepSeek-V4-Pro | 15.1 |