Tag: speech recognition

An open model sorts eight overlapping speakers in real time

NVIDIA released Nemotron 3 Diarization, a small open model that tops VoiceArena's leaderboard while running live conversation streams.

SpaceXAI doubles speech accuracy while holding its transcribe price

Grok Voice Transcribe 2.0 keeps batch pricing at $0.10 per hour of audio and claims the top accuracy…

OpenAI prices full-duplex voice at 5 cents a minute

GPT-Live-1 reaches the API, letting developers put natural turn-taking speech into products at a per-minute price.

Microsoft claims a 60-language speech record for MAI-Transcribe-2

The model posts a 5.2% average error rate on FLEURS and prices at $0.10 per hour of audio.

Meta folds transcription, speaker ID and silence detection into one model

Muse Voice Transcribe handles streaming speech-to-text, 20-plus speaker diarization, and endpointing in a single pass.

Gemini 3.5 Transcribe trims Google speech errors to 2.6 percent

Google's new speech model posts a 2.6 percent average error rate across 85-plus languages.