2 September 2026
Meta releases Muse Voice Transcribe audio model
First reported
TLDR AI and The Rundown AI ran this on , all on the same day.
- Meta built Muse Voice Transcribe to convert speech to text in real time while the person is still talking.
- The model can identify and separate the voices of 20 or more different speakers in the same audio, useful for meeting transcripts.
- It handles code-switching (switching between languages mid-sentence) and can use context clues to improve accuracy.
Where they differ
TLDR AIfocused on technical features like endpointing and contextual biasing that most readers would not encounter.
The Rundown AIemphasized that the model ranks first on a transcription leaderboard.
What each one reported
TLDR AITLDR editorial team
Muse Voice Transcribe is Meta's first real-time audio perception model supporting streaming speech recognition, diarization for 20+ speakers, endpointing, multilingual code-switching, and contextual biasing.
The Rundown AIRowan Cheung
Meta released Muse Voice Transcribe, a real-time speech-to-text model that identifies and differentiates 20 or more speakers simultaneously and tops Artificial Analysis' leaderboard for voice transcription.