UI4A-Bench
Code · 2026-07-18
UI4A-Bench evaluates models on 161 cases across eight experience domains by asking them to generate code-native interactive interfaces from natural-language needs without schema guidance. More than 200 atomic rubrics cover delivery, text fidelity, visual quality, and interaction correctness; simulator-tested interactions and Gemini 3.5 Flash rubric judgments are combined with reliability-aware weighting and normalized to 100 points.
Top models (higher is better)
| Model | Score |
|---|---|
| Macaron-V1-Venti | 87.8 |
| Opus 4.8 | 75.9 |
| GPT-5.5 | 72.1 |
| GLM-5.2 | 67.1 |
| MiniMax M3 | 63.0 |
| Qwen3.7-Max | 62.5 |
| Gemini 3.1 Pro Preview | 60.3 |