NVIDIA released Nemotron 3 Diarization, a small open model that tops VoiceArena's leaderboard while running live conversation streams.
Grok Voice Transcribe 2.0 keeps batch pricing at $0.10 per hour of audio and claims the top accuracy…
GPT-Live-1 reaches the API, letting developers put natural turn-taking speech into products at a per-minute price.
The model posts a 5.2% average error rate on FLEURS and prices at $0.10 per hour of audio.
Muse Voice Transcribe handles streaming speech-to-text, 20-plus speaker diarization, and endpointing in a single pass.
Google's new speech model posts a 2.6 percent average error rate across 85-plus languages.