Voxtral Transcribes at the Speed of Sound: Introducing Voxtral Transcribe 2
Today, we're excited to unveil Voxtral Transcribe 2, a groundbreaking leap in speech-to-text technology. This release introduces two cutting-edge models: Voxtral Mini Transcribe V2 for batch transcription and Voxtral Realtime for real-time applications. Voxtral Realtime is open-source under the Apache 2.0 license, offering unparalleled flexibility and control.
We're also launching an interactive audio playground in Mistral Studio, allowing you to test transcription instantly with diarization and timestamps. Here's a breakdown of what makes Voxtral Transcribe 2 a game-changer:
Voxtral Mini Transcribe V2: The Ultimate Transcription Powerhouse
- State-of-the-Art Transcription: Achieves industry-leading accuracy with speaker diarization, context biasing, and word-level timestamps in 13 languages.
- Speaker Diarization: Accurately identifies speakers and provides precise start/end times, perfect for meeting transcription, interview analysis, and multi-party calls.
- Context Biasing: Guides the model towards correct spellings of names, technical terms, and domain-specific vocabulary, ensuring accuracy for proper nouns and industry jargon.
- Word-Level Timestamps: Enables subtitle generation, audio search, and content alignment with precise timestamps for each word.
- Expanded Language Support: Supports 13 languages, with non-English performance significantly outperforming competitors.
- Noise Robustness: Maintains accuracy in challenging acoustic environments, from factory floors to busy call centers.
- Longer Audio Support: Processes recordings up to 3 hours in a single request.
Voxtral Realtime: Real-Time Transcription at Sub-200ms Latency
- Purpose-Built for Live Transcription: Designed for applications where latency is critical, achieving sub-200ms delay.
- Multilingual Excellence: Transcribes audio in 13 languages with near-offline accuracy, making it ideal for voice agents and real-time applications.
- Industry-Leading Efficiency: Offers the best price-performance ratio, outperforming competitors like GPT-4o mini Transcribe, Gemini 2.5 Flash, Assembly Universal, and Deepgram Nova.
- Open-Source: Available as open weights under Apache 2.0, allowing for deployment on edge devices for privacy-first applications.
Transforming Voice Applications
Voxtral empowers a wide range of voice applications across diverse industries:
- Meeting Intelligence: Transcribe multilingual recordings with speaker diarization, making it easy to attribute who said what and when. Ideal for annotating large volumes of meeting content at industry-leading cost efficiency.
- Voice Agents and Virtual Assistants: Build conversational AI with sub-200ms transcription latency, creating natural and responsive voice interfaces.
- Contact Center Automation: Transcribe calls in real-time, enabling AI systems to analyze sentiment, suggest responses, and populate CRM fields while conversations are happening.
- Media and Broadcast: Generate live multilingual subtitles with minimal latency, handling proper nouns and technical terminology with ease.
- Compliance and Documentation: Monitor and transcribe interactions for regulatory compliance, ensuring clear speaker attribution and precise audit trails.
Get Started
Voxtral Mini Transcribe V2 is available now via API at $0.003 per minute. Try it in the Mistral Studio audio playground or Le Chat. Voxtral Realtime is available via API at $0.006 per minute and as open weights on Hugging Face.
Explore our documentation on audio and transcription capabilities, and join us in shaping the future of speech AI.
We're Hiring!
If you're passionate about building world-class speech AI and empowering developers worldwide, we want to hear from you! Apply to join our team and be part of this exciting journey.