Atlas

Benchmarks

← All benchmarks

MRCR v2 (8-needle) - 1M pointwise

General QA · 2026-03-03

This MRCR v2 8-needle metric reports pointwise string-similarity performance at a one-million-token context length rather than a cumulative average. It tests whether a model can reproduce the correct instance from an adversarially similar multi-round conversation at its full context limit.

Top models (higher is better)

ModelScore
Gemini 3.6 Flash54.0
Gemini 3.5 Flash26.6
Gemini 3.1 Pro Preview26.3
Gemini 3 Pro Preview26.3
Gemini 3 Flash Preview22.1
Gemini 3.5 Flash-Lite21.3
Gemini 2.5 Flash21.0
Gemini 2.5 Pro16.4
Gemini 2.5 Flash Preview (09-2025)16.3
Gemini 3.1 Flash-Lite Preview12.3
Gemini 3.1 Flash-Lite12.3
Gemini 2.5 Flash-Lite Preview (09-2025)7.7
Grok 4.1 Fast6.1
Gemini 2.5 Flash-Lite5.4
Gemini 2.0 Flash5.3
Loading Atlas data…