BBH (Open LLM Leaderboard v2)
General QA · 2024-06-26
BIG-Bench Hard collects tasks that were difficult for contemporary models at release. The Open LLM Leaderboard v2 figure averages accuracy over a subset of BBH subtasks whose option counts differ, so neither a single denominator nor one guessing floor is well defined. Scored by HuggingFace's Open LLM Leaderboard v2 under lm-evaluation-harness with a fixed prompt template and no chain-of-thought, which is a different measurement protocol from lab-reported and Artificial Analysis runs of the same underlying benchmark, so it forms its own factor group rather than pooling with them. Raw accuracy is imported, not the leaderboard's baseline-normalized column.