2 September 2026

Meta releases Muse Voice Transcribe audio model

First reported

TLDR AI and The Rundown AI ran this on , all on the same day.

  • Meta built Muse Voice Transcribe to convert speech to text in real time while the person is still talking.
  • The model can identify and separate the voices of 20 or more different speakers in the same audio, useful for meeting transcripts.
  • It handles code-switching (switching between languages mid-sentence) and can use context clues to improve accuracy.

Where they differ

  • TLDR AI

    focused on technical features like endpointing and contextual biasing that most readers would not encounter.

  • The Rundown AI

    emphasized that the model ranks first on a transcription leaderboard.

What each one reported

TLDR AITLDR editorial team

Muse Voice Transcribe is Meta's first real-time audio perception model supporting streaming speech recognition, diarization for 20+ speakers, endpointing, multilingual code-switching, and contextual biasing.

The Rundown AIRowan Cheung

Meta released Muse Voice Transcribe, a real-time speech-to-text model that identifies and differentiates 20 or more speakers simultaneously and tops Artificial Analysis' leaderboard for voice transcription.