Atlas

Benchmarks

← All benchmarks

UI4A-Bench

Code · 2026-07-18

UI4A-Bench evaluates models on 161 cases across eight experience domains by asking them to generate code-native interactive interfaces from natural-language needs without schema guidance. More than 200 atomic rubrics cover delivery, text fidelity, visual quality, and interaction correctness; simulator-tested interactions and Gemini 3.5 Flash rubric judgments are combined with reliability-aware weighting and normalized to 100 points.

Top models (higher is better)

ModelScore
Macaron-V1-Venti87.8
Opus 4.875.9
GPT-5.572.1
GLM-5.267.1
MiniMax M363.0
Qwen3.7-Max62.5
Gemini 3.1 Pro Preview60.3
Loading Atlas data…