Voxtral Small 24B is Mistral's lightweight multimodal model that processes both text and audio with a 32K token context window.
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.050 | $0.150 | - | |
| $0.100 | $0.300 | - | |
| $0.100 | $0.300 | $0.010 |
Prices updated daily. Last check: Sep 6, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Voxtral Small 24B is well-suited for applications requiring efficient audio and text processing at scale. Common use cases include audio transcription services, voice-to-text applications, podcast analysis, customer service call processing, and content moderation for audio platforms. Its lightweight architecture makes it appropriate for high-volume scenarios where audio understanding is needed but computational budgets are constrained. The model works well for building voice interfaces, analyzing recorded meetings, and processing multimedia content where both spoken and written elements need to be understood together.
Voxtral Small 24B pricing varies by provider and usage patterns. As a lightweight multimodal model, costs will differ for text versus audio processing. Check the pricing table above for current rates across all providers offering this model.
Voxtral Small 24B excels at audio transcription, voice analysis, and applications requiring both text and audio understanding. Its lightweight design makes it ideal for high-volume scenarios like customer service call analysis, podcast processing, and voice interface development where efficiency is important.
No, Voxtral Small 24B does not support tool calling or function execution capabilities. It focuses specifically on text and audio processing tasks. If you need function calling alongside multimodal capabilities, you would need to consider other models in Mistral's lineup or alternative providers.