CAISI PortBench
Agents · 2026-05-01
CAISI PortBench is a non-public software-engineering benchmark that asks agents to port command-line tools between programming languages from a reference implementation. Scores report the maximum percentage of hidden tests passed by the ported implementation.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Mythos Preview | 80.1 |
| GPT-5.5 | 78.0 |
| Opus 4.8 | 60.7 |
| Opus 4.6 | 60.0 |
| GLM-5.2 | 41.7 |
| GPT-5.4 Mini | 41.0 |