Atlas

Benchmarks

← All benchmarks

Audio MultiChallenge (Audio Output)

Multimodal · 2025-12-17

Audio MultiChallenge evaluates multi-turn conversational intelligence for audio-language models using 452 natural spoken conversations. This audio-output track measures the percentage of conversations whose final native-speech response satisfies every atomic rubric; the accompanying text stream is evaluated under the same rubrics.

Top models (higher is better)

ModelScore
Qwen3 Omni 30B A3B Instruct24.3
Loading Atlas data…