Round-Trip Translation Mean Score
Chat & Writing · 2025-09-15
Lech Mazur's Round-Trip Translation benchmark evaluates how much meaning and voice survive translation out of English and back to English. The headline leaderboard reports an ensemble mean score.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5 | 8.7 |
| Grok 4 | 8.6 |
| Opus 4.1 | 8.6 |
| Gemini 2.5 Pro | 8.5 |
| Qwen3 Max (rolling alias) | 8.3 |
| DeepSeek-V3.1 | 8.3 |
| Mistral Medium 3.1 | 8.3 |
| Kimi K2 Instruct 0905 | 8.3 |