Atlas

Benchmarks

← All benchmarks

HieroglyphBench

Multimodal · 2026-06-24

HieroglyphBench is a vision-language OCR benchmark for ancient Egyptian hieroglyphs. It contains 30 curated photographed inscription columns from the Pyramid of Unas, and models must output the ordered sequence of Gardiner sign-list codes; scores are sign accuracy from edit distance against expert transcriptions.

Top models (higher is better)

ModelScore
Gemini 3.5 Flash52.5
Gemini 3 Flash Preview48.8
Gemini 3.1 Pro Preview39.7
GPT-5.6 Sol27.6
Kimi K325.9
Muse Spark 1.124.5
Claude Fable 522.9
GPT-5.522.8
Grok 4.521.2
GPT-5.6 Terra20.5
Kimi K2.619.2
Gemini 2.5 Pro18.8
Grok 4.2016.8
GPT-5.4 Mini15.7
Opus 4.814.0
Inkling12.5
MiniMax M311.1
Qwen3.7-Plus10.1
GPT-4o9.5
Loading Atlas data…