Atlas

Benchmarks

← All benchmarks

OSWorld

Agents · 2024-04-11

A computer-use benchmark where agents complete real desktop and web tasks in reproducible OS environments using keyboard/mouse actions and structured UI observations.

Top models (higher is better)

ModelScore
Claude Mythos Preview79.6
GPT-5.475.0
Opus 4.672.7
Sonnet 4.672.1
Opus 4.566.3
o345.2
Sonnet 443.9
Gemini 2.5 Pro Preview 03-2541.4
Computer Use Preview38.1
Claude 3.7 Sonnet35.8
GPT-4o27.0
UI-TARS 72B DPO24.6
Claude 3.5 Sonnet (Oct 2024)22.0
Qwen2.5-VL-72B-Instruct8.8
Kimi-VL-A3B-Instruct8.2
Gemini 1.0 Pro Vision5.8
GPT-4 Turbo5.4
Gemini 1.5 Pro5.4
GPT-4o Mini3.8
Opus 32.4
LLaVA-OneVision 72B2.4
Qwen VL Max 08092.4
MiniCPM-V 2.61.9
CogAgent1.1
Loading Atlas data…