Atlas

Benchmarks

← All benchmarks

Breakpoint (CAISI 200-task subset)

Agents · 2025-09-30

CAISI's Breakpoint evaluation uses a fixed 200-task subset generated from the public remove-data task pool. Agents operate in CAISI's ReAct-style software-engineering environment, and a task succeeds only when the repaired repository passes all original tests.

Top models (higher is better)

ModelScore
GPT-598.0
gpt-oss-120b93.0
Opus 492.3
DeepSeek-V3.178.5
DeepSeek-R1-052860.2
DeepSeek-R116.0
Loading Atlas data…