Meta Superintelligence Labs is selling a speech model that merges three jobs usually split across separate systems. Muse Voice Transcribe performs streaming transcription, tracks more than 20 speakers at once, and detects when someone has stopped talking, all in a single autoregressive pass.
Most production voice stacks chain a transcriber, a speaker separator, and a voice-activity detector together, and each hand-off adds latency and a new failure mode. Meta positions Muse Voice Transcribe as its first real-time audio perception model, with no post-processing needed.
Pricing starts at $3.00 per 1,000 audio minutes, or $0.18 an hour, for the hosted service, which Meta lists on its Model API as muse-voice-transcribe-1.0. The model already handles dictation chores inside Meta AI for Mac and in Muse Code, according to Meta.
Voice agents, meeting transcription, and other real-time use cases are the obvious targets. The release also stretches Meta’s Muse family beyond image generation and coding into speech perception, deepening its bet on models that hear and act live.