OmniDocBench v1.5
Multimodal · 2025-09-25
OmniDocBench v1.5 evaluates end-to-end parsing of 1,355 diverse PDF pages. Its Overall score is the mean of text accuracy derived from normalized edit distance, table TEDS, and formula CDM.
Top models (higher is better)
| Model | Score |
|---|---|
| Kimi K3 | 91.1 |
| Gemini 3 Flash Preview | 90.4 |
| Gemini 3 Pro Preview | 90.2 |
| Claude Fable 5 | 89.8 |
| GPT-5.5 | 89.4 |
| Kimi K2.5 | 89.3 |
| Qwen3-VL-235B-A22B-Instruct | 89.2 |
| Gemini 2.5 Pro | 88.0 |
| Opus 4.8 | 87.9 |
| Qwen2.5-VL-72B-Instruct | 87.0 |
| GPT-5.6 Sol | 85.8 |
| GPT-5.2 (2025-12-11) | 85.8 |
| InternVL3-78B | 80.3 |
| GPT-4o (2024-08-06) | 75.0 |