Meta has launched Muse Voice Transcribe, a real-time speech recognition model that transcribes incoming audio and identifies speakers as a conversation unfolds. The release comes from Meta Superintelligence Labs and will reach developers through the Meta Model API, as first reported by VentureBeat.
Meta lists processing at $0.18 per hour of audio. It is also integrating the system into Meta AI for Mac and Muse Code, extending the technology beyond a standalone developer API.
Muse handles live audio in 80-millisecond segments
The service processes incoming sound in 80-millisecond chunks, allowing it to return text while a person is still speaking. That separates real-time transcription from batch systems, which produce a transcript only after receiving the full recording.
Muse supports audio sessions longer than one hour and can identify more than 20 speakers during transcription. Its feature set includes endpoint detection, which determines when speech starts and stops, along with speaker diarization, language biasing and keyword biasing.
- Streaming transcription for live audio
- Speaker identification for conversations with more than 20 participants
- Multilingual recognition, including speech that switches languages during a session
- Training across more than 70 languages, with 25 extensively validated for the initial release
Speaker identification is central to enterprise use cases
Diarization assigns each transcript segment to a specific participant, making it useful for meeting records, call analysis, compliance workflows and voice agents. Meta’s stated threshold places Muse near Amazon Transcribe’s documented 30-speaker streaming limit, though Speechmatics says its real-time product supports 50 speakers by default and up to 100 with a higher setting.
The pricing could matter for organizations processing large volumes of calls or meetings, where per-hour costs accumulate quickly. Language switching support also targets workplaces and customer-service operations where a single conversation can move between languages without warning.
Meta has not specified the launch date or geographic availability for every access channel. Independent tests have yet to establish the model’s accuracy, latency or reliability, and Meta has not said whether its $0.18 hourly price applies equally across languages, usage volumes and API features. The company also has not published a maximum audio duration or an exact upper speaker limit beyond more than 20.
This article was produced with AI assistance from multi-source reporting and is published under our editorial standards.