Voxtral Small 1.0 24B 2507
Mistral AI · 2025-07-15 · 24.3B parameters
Voxtral Small 1.0 24B 2507 is Mistral AI's Apache-2.0 open-weight audio-language model built on the Mistral Small 3.1 24B text backbone. It combines a 32-layer Whisper-large-v3-derived audio encoder, a 4x-downsampling MLP adapter, and a 40-layer text decoder; the original BF16 checkpoint contains 24,259,880,960 learned parameters and supports an exact 32,768-token context. It accepts audio and text and generates text for transcription, translation, audio understanding, summarization, function calling, and text-only use.
Benchmark scores
| Benchmark | Score |
|---|---|
| ObviousBench | 43.8 |