Atlas

Benchmarks

← All benchmarks

Phare Tools Reliability (French)

Safety · 2025-03-27

Phare task testing reliability when using or reasoning about tool/plugin outputs. This row is the French language split.

Top models (higher is better)

ModelScore
Sonnet 4.698.8
Opus 4.697.7
Gemini 3.1 Pro Preview96.4
Qwen3.7 Max Preview96.2
Sonnet 595.3
DeepSeek-V4-Flash95.3
DeepSeek-V4-Pro95.2
Gemini 3.5 Flash94.9
Sonnet 4.594.1
GLM-5.294.1
Opus 4.593.3
GPT-5.593.2
Claude 3.5 Sonnet (Oct 2024)92.7
Qwen3.7-Plus92.7
Grok 492.6
Opus 4.192.0
Gemini 3 Pro Preview90.9
Kimi K2.590.8
Gemma 4 31B IT90.1
Mistral Medium 3.190.0
Mistral Medium 3.589.6
Haiku 4.588.9
Grok 388.9
Grok 4.388.8
Mistral Large 2 (Instruct 2407)88.6
GPT-5.188.1
Haiku 3.588.0
Mistral Large 3 675B Instruct 251287.5
Gemini 3.1 Flash-Lite87.3
GPT-4.1 Mini86.7
Mistral Small 3.1 24B Instruct 250386.6
GPT-4.186.3
GPT-5 Mini85.3
GPT-4o85.1
GPT-5.284.9
GPT-5 Nano84.9
gpt-oss-120b84.0
Kimi K2.684.0
Claude 3.7 Sonnet83.7
Grok 4 Fast83.0
Loading Atlas data…