Atlas

Benchmarks

← All benchmarks

CAISI PortBench

Agents · 2026-05-01

CAISI PortBench is a non-public software-engineering benchmark that asks agents to port command-line tools between programming languages from a reference implementation. Scores report the maximum percentage of hidden tests passed by the ported implementation.

Top models (higher is better)

ModelScore
Claude Mythos Preview80.1
GPT-5.578.0
Opus 4.860.7
Opus 4.660.0
GLM-5.241.7
GPT-5.4 Mini41.0
Loading Atlas data…