Atlas

Benchmarks

← All benchmarks

Surface Evolver Bench

Agents · 2026-06-04

Surface Evolver Bench evaluates agentic LLMs writing Surface Evolver .fe simulation datafiles for liquid-surface physics. Models can read curated documentation, run candidate files in Surface Evolver, repair them, and submit a structured datafile; the overall score combines static file checks, hidden Evolver execution, and hidden metric checks against reference solutions.

Top models (higher is better)

ModelScore
Claude Fable 595.0
Kimi K395.0
GPT-5.6 Sol93.1
GPT-5.588.1
Opus 4.887.5
GPT-5.6 Terra83.8
Grok 4.574.4
Sonnet 560.0
GPT-5.6 Luna60.0
Gemini 3.5 Flash58.1
GLM-5.255.6
MiniMax M353.1
Muse Spark 1.152.5
Kimi K2.7 Code48.8
Qwen3.6 35B A3B44.4
DeepSeek-V4-Pro40.0
Gemma 4 31B IT30.6
gpt-oss-120b25.0
Laguna M.115.6
Trinity Large Thinking15.6
Loading Atlas data…