Atlas

Benchmarks

← All benchmarks

IFBench (Artificial Analysis)

Chat & Writing · 2025-07-02

IFBench evaluates precise instruction following on the 294-question Artificial Analysis implementation of AllenAI's 58 diverse, verifiable out-of-domain constraints. This evaluator-specific component measures whether models generalize to new output rules and formatting demands; higher scores mean more constraints satisfied.

Top models (higher is better)

ModelScore
Inkling-Small Preview83.4
Grok 4.383.3
Grok 4.2082.9
MiniMax M382.9
Inkling-Small82.2
Nemotron 3 Ultra 550B A55B81.4
Qwen3.7-Max80.5
Nemotron Cascade 2 30B A3B80.4
MiMo-V2.5-Pro79.9
Inkling79.8
Nova 2 Pro (Preview)79.6
DeepSeek-V4-Flash79.2
Qwen3.5 397B A17B78.8
Gemini 3.5 Flash-Lite78.6
Gemini 3 Flash Preview78.0
Qwen3.7-Plus78.0
GPT-5.2-Codex77.6
Gemini 3.1 Flash-Lite Preview77.2
Gemini 3.1 Pro Preview77.1
Qwen3.6-Max-Preview76.6
DeepSeek-V4-Pro76.5
Gemini 3.5 Flash76.3
GLM-5.176.3
Kimi K2.676.0
GPT-5.4 Nano75.9
Muse Spark75.9
GPT-5.575.9
MiniMax M2.775.7
Qwen3.5 122B A10B75.7
Gemma 4 31B IT75.6
Qwen3.5-27B75.6
GPT-5.275.4
GPT-5 Mini75.4
GPT-5.3-Codex75.4
Qwen3.6 Plus (2026-04-02)75.2
GPT-5-Codex74.1
Command A+ (May 2026, BF16)73.9
GPT-5.473.9
Gemma 4 12B IT73.5
GLM-5.273.3
Loading Atlas data…