IFEval (Open LLM Leaderboard v2)
Chat & Writing · 2024-06-26
IFEval measures whether a model obeys verifiable instruction constraints such as length, format, and forbidden words. The Open LLM Leaderboard v2 figure averages prompt-level and instruction-level strict accuracy, so it is a mean of two rates rather than a single item proportion. Scored by HuggingFace's Open LLM Leaderboard v2 under lm-evaluation-harness with a fixed prompt template and no chain-of-thought, which is a different measurement protocol from lab-reported and Artificial Analysis runs of the same underlying benchmark, so it forms its own factor group rather than pooling with them. Raw accuracy is imported, not the leaderboard's baseline-normalized column.