Kyutai released two open-weight speech-to-speech models that reason through math without transcribing audio first, posting a large jump…
NVIDIA released Nemotron 3 Diarization, a small open model that tops VoiceArena's leaderboard while running live conversation streams.
Gander keeps a conversation alive in one model while a second model does the work.
Grok Voice Transcribe 2.0 keeps batch pricing at $0.10 per hour of audio and claims the top accuracy…
Reykjavik-based Treble has extended its Series A to build acoustic simulation infrastructure for voice models, wearables and robots.
The two new audio models switch between 97 languages mid-sentence and run tool calls in the background while…
GPT-Live-1 reaches the API, letting developers put natural turn-taking speech into products at a per-minute price.
Muse Voice Transcribe handles streaming speech-to-text, 20-plus speaker diarization, and endpointing in a single pass.
A September 2 vote clears the final hurdle and puts the conversational AI merger on track to close…
Plaud's new earbuds record and process conversations through a 4G charging case that works without a phone.
Google's new speech model posts a 2.6 percent average error rate across 85-plus languages.
Cartesia's Sonic-3.6 tops both Artificial Analysis speech leaderboards, winning the controlled-voice test that isolates engine quality.