Atlas

Benchmarks

← All benchmarks

Spider 2.0-Lite (SQLite subset)

Code · 2026-05-11

The 135-case SQLite-backed subset of Spider 2.0-Lite used by Interfaze's public runner. It evaluates text-to-SQL execution accuracy on the local examples and is distinct from the full 547-task Spider 2.0-Lite benchmark.

Top models (higher is better)

ModelScore
Interfaze Beta52.9
Sonnet 550.6
Gemini 3.5 Flash46.7
Grok 4.345.9
Gemini 3 Flash Preview45.2
GPT-5.4 Mini26.7
Loading Atlas data…