Surface Evolver Bench
Agents · 2026-06-04
Surface Evolver Bench evaluates agentic LLMs writing Surface Evolver .fe simulation datafiles for liquid-surface physics. Models can read curated documentation, run candidate files in Surface Evolver, repair them, and submit a structured datafile; the overall score combines static file checks, hidden Evolver execution, and hidden metric checks against reference solutions.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Fable 5 | 95.0 |
| Kimi K3 | 95.0 |
| GPT-5.6 Sol | 93.1 |
| GPT-5.5 | 88.1 |
| Opus 4.8 | 87.5 |
| GPT-5.6 Terra | 83.8 |
| Grok 4.5 | 74.4 |
| Sonnet 5 | 60.0 |
| GPT-5.6 Luna | 60.0 |
| Gemini 3.5 Flash | 58.1 |
| GLM-5.2 | 55.6 |
| MiniMax M3 | 53.1 |
| Muse Spark 1.1 | 52.5 |
| Kimi K2.7 Code | 48.8 |
| Qwen3.6 35B A3B | 44.4 |
| DeepSeek-V4-Pro | 40.0 |
| Gemma 4 31B IT | 30.6 |
| gpt-oss-120b | 25.0 |
| Laguna M.1 | 15.6 |
| Trinity Large Thinking | 15.6 |