Audio MultiChallenge (Audio Output)
Multimodal · 2025-12-17
Audio MultiChallenge evaluates multi-turn conversational intelligence for audio-language models using 452 natural spoken conversations. This audio-output track measures the percentage of conversations whose final native-speech response satisfies every atomic rubric; the accompanying text stream is evaluated under the same rubrics.
Top models (higher is better)
| Model | Score |
|---|---|
| Qwen3 Omni 30B A3B Instruct | 24.3 |