BigCodeBench-Hard Complete
Code · 2024-06-18
BigCodeBench evaluates code generation on practical programming tasks that compose calls across many libraries, scored by test execution as Pass@1. BigCodeBench-Hard is the curated 148-task subset that the maintainers select as most demanding; this is its Complete split. The 148-task denominator is confirmed by the reported score lattice. Complete and Instruct are two prompt protocols over the same task pool, and the Hard subset is nested inside the full set, so all four scores share one factor group: they measure one construct and would otherwise supply four correlated votes for it.