Two new spoken dialogue models from Google landed on September 15, and they are built for opposite jobs.
One, Gemini 3.8 Live, is cheap and fast, meant for the kind of back-and-forth a support bot has with a caller. The other, Gemini 3.8 Live Extended Thinking, is slow on purpose. It works through several steps of reasoning and talks while it does, so a user hears an acknowledgement right away and then a running account of what it is doing.
Both are live now. Developers reach them through the Gemini API and AI Studio. Enterprises route through Gemini Enterprise. Consumers meet them in the Gemini app, in Search Live and inside Workspace.
Tom Ouyang, a principal engineer, and Malini Jaganathan of the technical staff signed the post for the Gemini Audio Team.
Google leaned on outside scoreboards for evidence. Artificial Analysis put the thinking model first on its Speech to Speech Quality Index at 82.6. The same model scored 68.6 percent on tau-Voice and 35.1 percent on Sierra’s banking version of that test, plus 97.7 percent on Big Bench Audio. Its faster sibling came second in the Speech Agent Arena. ServiceNow’s EVA-Bench, run on the Live API inside Gemini Enterprise, was cited to show both models holding accuracy and conversational quality together.
For teams shipping voice agents, two details stand out. Language switching happens mid-conversation across 97 options with no restart. And tool or API calls run in the background while the person keeps talking, which removes the silent gap that usually follows a request.