Breakpoint (CAISI 200-task subset)
Agents · 2025-09-30
CAISI's Breakpoint evaluation uses a fixed 200-task subset generated from the public remove-data task pool. Agents operate in CAISI's ReAct-style software-engineering environment, and a task succeeds only when the repaired repository passes all original tests.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5 | 98.0 |
| gpt-oss-120b | 93.0 |
| Opus 4 | 92.3 |
| DeepSeek-V3.1 | 78.5 |
| DeepSeek-R1-0528 | 60.2 |
| DeepSeek-R1 | 16.0 |