Atlas

Benchmarks

← All benchmarks

Round-Trip Translation Mean Score

Chat & Writing · 2025-09-15

Lech Mazur's Round-Trip Translation benchmark evaluates how much meaning and voice survive translation out of English and back to English. The headline leaderboard reports an ensemble mean score.

Top models (higher is better)

ModelScore
GPT-58.7
Grok 48.6
Opus 4.18.6
Gemini 2.5 Pro8.5
Qwen3 Max (rolling alias)8.3
DeepSeek-V3.18.3
Mistral Medium 3.18.3
Kimi K2 Instruct 09058.3
Loading Atlas data…