Atlas

Benchmarks

← All benchmarks

MedQA Demographic Bias (Vals) - Black

Science · 2026-04-16

The black condition of Vals AI's MedQA demographic-framing evaluation. The source does not expose a stable per-run denominator, so it is kept separate from standard MedQA.

Top models (higher is better)

ModelScore
Gemini 3.1 Pro Preview96.8
O196.6
GPT-5.496.5
GPT-596.5
GPT-5 Mini96.3
GPT-5.196.2
O4 Mini96.1
Opus 4.596.0
Gemini 3 Pro Preview95.8
o395.8
Gemini 3 Flash Preview95.7
Qwen3.5 Plus (2026-02-15)95.7
Opus 4.695.3
GPT-5.294.8
o3-mini94.8
Sonnet 4.594.7
Grok 4.2094.5
GLM-594.4
GLM-4.794.3
Kimi K2.594.2
DeepSeek-V3.294.0
Opus 4.193.8
o1 Preview93.6
GPT-5 Nano93.2
Gemini 2.5 Pro Experimental 03-2593.1
Kimi K2 Thinking92.8
Sonnet 4.692.8
Sonnet 492.5
Grok 4.1 Fast92.3
Opus 492.3
Grok 292.2
MiniMax M2.592.1
Grok 4 Fast91.9
Grok 491.8
GLM 4.691.8
DeepSeek-R191.2
MiniMax M2.191.2
Gemini 2.5 Flash Preview (09-2025)90.9
gpt-oss-120b90.8
GPT-4.190.8
Loading Atlas data…