Atlas

Benchmarks

← All benchmarks

RadLE 2.0 Reliability Index (RadLE-R)

Multimodal

RadLE-R measures diagnostic accuracy among answers submitted with autonomous-level confidence (Likert 3 or 4). It tests whether the answers a system presents most confidently can be trusted.

Top models (higher is better)

ModelScore
Human Expert Baseline (RadLE 2.0)55.0
Claude Fable 554.7
Muse Spark 1.140.2
Gemini 3.1 Pro Preview32.5
GPT-5.6 Sol Pro31.9
Qwen3.7-Plus18.7
GLM 5V Turbo17.2
Grok 4.515.5
Gemma 4 31B (RadLE 2.0 label)9.4
Llama 4 Maverick (RadLE 2.0 label)7.9
MiniMax M37.3
OctoMed-7B3.5
Lingshu-32B3.5
Nemotron 3 Nano Omni 30B A3B0.9
MedGemma 1.5 4B IT0.7
InternVL3.5-8B0.6
Mistral Large 3 2512 (RadLE 2.0 label)0.5
Loading Atlas data…