BuseyBench SVG
Multimodal · 2026-07-02
BuseyBench SVG measures how well a model creates a standalone 1024×1024 SVG portrait of Gary Busey from one fixed no-web prompt. Three vision judges each score the rendered output three times; BuseyBench takes each judge's median and then averages the judges. Its raw 0–10 score combines Busey likeness (50%), face coherence (20%), visual aesthetics (15%), and prompt adherence (15%); Atlas stores this raw absolute score rather than the leaderboard's pairwise-adjusted ranking score.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.6 Terra Pro | 6.5 |
| GPT-5.6 Sol Pro | 6.3 |
| Claude Opus 5 | 6.3 |
| Gemini 3.1 Pro Preview | 6.3 |
| Kimi K3 | 6.2 |
| GPT-5.6 Sol | 6.1 |
| GPT-5.6 Terra | 6.1 |
| GPT-5.5 Pro | 6.1 |
| GPT-5.5 | 6.0 |
| Gemini 3.5 Flash | 5.9 |
| Claude Fable 5 | 5.9 |
| Fugu Ultra v1.0 | 5.8 |
| GPT-5.6 Luna Pro | 5.7 |
| Opus 4.6 | 5.5 |
| GPT-5.2 Pro | 5.5 |
| Nex-N2-Pro | 5.5 |
| Opus 4.8 | 5.4 |
| Kimi K2.6 | 5.2 |
| Opus 4.5 | 5.2 |
| GPT-5.2-Codex | 5.2 |
| Qwen3.7-Plus | 5.1 |
| Opus 4.7 | 5.1 |
| GLM 5V Turbo | 5.1 |
| GPT-5.2 | 5.1 |
| Gemini 3.6 Flash | 5.0 |
| GLM-5.2 | 5.0 |
| Qwen3.5 Plus (2026-02-15) | 4.9 |
| Grok 4.5 | 4.9 |
| GPT-5.1-Codex-Max | 4.9 |
| GPT-5.3-Codex | 4.9 |
| Sonnet 4.6 | 4.8 |
| GPT-5.3 Instant | 4.8 |
| Sonar Pro | 4.7 |
| Muse Spark 1.1 | 4.7 |
| Qwen3.7-Max | 4.6 |
| o3 | 4.6 |
| GPT-5.4 | 4.6 |
| Sonnet 5 | 4.5 |
| GPT-5 | 4.5 |
| Qwen3.6 27B | 4.5 |