Atlas

Benchmarks

← All benchmarks

Arena Text — Medicine & Healthcare (No Style Control)

Professional Work

Raw Arena rating for medicine and healthcare industry prompts without style-control adjustment. Higher ratings indicate stronger preference in blind pairwise user votes.

Top models (higher is better)

ModelScore
Claude Opus 51514
Opus 4.61506
Opus 4.71497
Gemini 3 Pro Preview1493
Kimi K31492
Gemini 3.5 Flash1490
ERNIE 5.11489
Gemini 3.1 Pro Preview1483
Claude Fable 51483
Muse Spark1483
GLM-5.21478
Gemini 3.6 Flash1478
Gemini 2.5 Pro1476
Qwen3.7 Max Preview1475
GLM-5.11474
MiMo-V2.5-Pro1472
Muse Spark 1.11469
DeepSeek-V3.2-Exp1468
Qwen3 Max Preview1468
DeepSeek-V3.1-Terminus1467
Qwen3.7-Plus1467
Gemini 3 Flash Preview1466
DeepSeek-V4-Pro1465
GLM-4.71464
LongCat-Flash-Chat1462
GPT-5.51461
Opus 4.81459
GPT-5.41459
GLM 4.61459
GPT-5.11457
Qwen3 Next 80B A3B Instruct1456
GLM-4.51454
Grok 4.11454
Qwen3.5 397B A17B1454
Kimi K2.61453
MiniMax M31453
GLM-51453
Kimi K2.51452
Sonnet 4.61452
Mistral Medium 3.11452
Loading Atlas data…