Atlas

Benchmarks

← All benchmarks

IFEval (Open LLM Leaderboard v2)

Chat & Writing · 2024-06-26

IFEval measures whether a model obeys verifiable instruction constraints such as length, format, and forbidden words. The Open LLM Leaderboard v2 figure averages prompt-level and instruction-level strict accuracy, so it is a mean of two rates rather than a single item proportion. Scored by HuggingFace's Open LLM Leaderboard v2 under lm-evaluation-harness with a fixed prompt template and no chain-of-thought, which is a different measurement protocol from lab-reported and Artificial Analysis runs of the same underlying benchmark, so it forms its own factor group rather than pooling with them. Raw accuracy is imported, not the leaderboard's baseline-normalized column.

Top models (higher is better)

ModelScore
Llama 3.3 70B Instruct90.0
Llama 3.1 70B Instruct86.7
Solar Pro Instruct Preview84.2
Mistral Large 2.1 (Instruct 2411)84.0
Qwen2.5 Instruct 32B83.5
Llama 3.1 Tülu 3 70B DPO82.8
Llama-3.1-Tülu-3-8B82.7
Qwen2.5 14B Instruct81.6
Qwen2 72B Instruct79.9
Gemma 2 27B IT79.8
Command R7B (Dec 2024)77.1
C4AI Command R+76.6
Hermes 3 Llama 3.1 70B76.6
Qwen2.5 7B Instruct75.9
C4AI Command R+ 08-202475.4
Gemma 2 9B IT74.4
Llama 3.2 3B Instruct73.9
Llama 3.1 Nemotron 70B Instruct HF73.8
Phi-4-mini-instruct73.8
Aya Expanse 32B73.0
Qwen2.5 Coder 32B Instruct72.7
Mixtral 8x22B Instruct71.8
Phi-3.5-MoE-instruct69.2
Command R67.5
Mistral Small Instruct 240966.7
Phi-3-Small-8K-Instruct65.0
Qwen2.5 3B Instruct64.7
Aya-23-35B64.6
Phi-3 Medium 4K Instruct64.2
Mistral NeMo Instruct 240763.8
Qwen2.5-Coder-7B-Instruct61.0
Yi 1.5 34B Chat60.7
Yi 1.5 9B Chat60.5
Phi-3 Medium 128K Instruct60.4
OpenChat 3.5 121060.4
Qwen2 VL 72B Instruct59.8
Phi-3 Mini 128K Instruct59.8
Qwen1.5 110B Chat59.4
Ministral 8B Instruct 241059.0
Phi-3.5 Mini Instruct57.7
Loading Atlas data…