Atlas

Benchmarks

← All benchmarks

MedQA Demographic Bias (Vals) - White

Science · 2026-04-16

The white condition of Vals AI's MedQA demographic-framing evaluation. The source does not expose a stable per-run denominator, so it is kept separate from standard MedQA.

Top models (higher is better)

ModelScore
O196.7
GPT-5.196.5
GPT-596.3
Gemini 3 Flash Preview96.3
Gemini 3.1 Pro Preview96.3
Opus 4.596.2
GPT-5 Mini96.2
Gemini 3 Pro Preview96.2
GPT-5.496.0
O4 Mini96.0
o396.0
Opus 4.695.8
Qwen3.5 Plus (2026-02-15)95.2
Sonnet 4.594.8
o3-mini94.7
Grok 4.2094.5
GLM-594.4
Kimi K2.594.3
DeepSeek-V3.294.1
GLM-4.794.1
GPT-5.294.0
Opus 4.194.0
GPT-5 Nano93.7
o1 Preview93.5
Opus 493.3
Gemini 2.5 Pro Experimental 03-2593.3
Grok 493.2
Sonnet 493.0
Kimi K2 Thinking92.8
Grok 4.1 Fast92.5
Grok 292.5
MiniMax M2.592.3
Gemini 2.5 Flash Preview (09-2025)92.3
GLM 4.692.2
Grok 4 Fast92.2
Sonnet 4.692.0
Gemini 2.5 Flash Preview 04-1791.9
GPT-4.191.5
gpt-oss-120b91.3
DeepSeek-R191.0
Loading Atlas data…