Google now sells synthetic speech at two speeds, and the difference is money rather than quality of intent.
The split is about economics as much as audio. Flash TTS is the expressive tier, where character work and creative direction matter most. Flash-Lite TTS serves high-volume production, trading warmth for lower latency and a smaller bill. Both reached the Gemini API and AI Studio this week.
Languages are the clearest split. Developers get 130 of them on the flagship model at launch. The Lite tier reaches 101. A shared library adds more than 2,000 ready-made voices on top.
Delivery can be scripted line by line. Instructions written into the text ask for a pause, a faster pace or a breath in the middle of a sentence, which removes the need for a separate markup layer.
Custom voices take a second path, built from a 30-second sample. Google requires the recording to belong to the speaker or be licensed, and a spoken consent clip must match the reference audio before the replica unlocks.
Provenance is attached rather than optional. Every clip carries an inaudible SynthID watermark and a C2PA record noting when it was created and whether it has been altered.
Self-hosting is off the table. No open weights were published, and enterprise access is listed as coming soon through Gemini Enterprise.