Atlas

Benchmarks

← All benchmarks

Phare Tools Reliability (Spanish)

Safety · 2025-03-27

Phare task testing reliability when using or reasoning about tool/plugin outputs. This row is the Spanish language split.

Top models (higher is better)

ModelScore
Sonnet 4.696.8
Opus 4.595.1
Opus 4.694.7
Gemini 3.1 Pro Preview94.7
DeepSeek-V4-Pro94.6
Sonnet 4.594.4
Gemini 3.5 Flash94.4
Sonnet 594.2
Opus 4.193.4
GLM-5.293.4
Gemini 3.1 Flash-Lite92.2
Grok 491.7
Kimi K2.591.0
DeepSeek-V4-Flash90.8
Qwen3.7 Max Preview90.5
Gemini 3 Pro Preview89.6
Qwen3.7-Plus88.9
Claude 3.5 Sonnet (Oct 2024)88.8
Gemma 4 31B IT88.8
GPT-5.587.9
Kimi K2.686.4
Haiku 4.586.0
Grok 4.384.5
Grok 283.2
GPT-4.182.3
Haiku 3.581.8
Mistral Large 2 (Instruct 2407)81.8
Mistral Medium 3.581.8
Grok 381.5
Mistral Large 3 675B Instruct 251280.4
GPT-4o Mini80.3
GPT-5.180.3
Claude 3.7 Sonnet79.9
GPT-4o79.9
GPT-5 Mini79.9
GPT-5.279.8
DeepSeek-V379.1
Mistral Medium 3.178.2
Mistral Small 3.2 24B Instruct 250678.2
gpt-oss-120b78.0
Loading Atlas data…