BigCodeBench Complete
Code · 2024-06-18
BigCodeBench evaluates code generation on practical programming tasks that compose calls across many libraries, scored by test execution as Pass@1. The Complete split gives the model a docstring-style function signature to finish over the full 1,140-task set. Complete and Instruct are two prompt protocols over the same task pool, and the Hard subset is nested inside the full set, so all four scores share one factor group: they measure one construct and would otherwise supply four correlated votes for it.