Voice agents have a new benchmark leader: Sonic-3.6, the latest streaming text-to-speech model from Cartesia, tops both Artificial Analysis leaderboards.
The headline number, 1,283 Elo on the Provider Voice board, is less telling than the second: 1,123 Elo on Controlled Voice, where every entrant is forced onto the same eight reference voices. That format strips away voice-catalog advantages, and Sonic-3.6 still beats Sonic-3.5 and ElevenLabs Eleven v3, a signal the engine itself improved.
Architecture is the differentiator. Sonic runs on state space models instead of transformers, and Cartesia claims sub-90-millisecond time-to-first-audio, with 100ms transcript latency on its Ink-2 speech recognition model.
It ships as a hosted API in beta with no open weights, at $49 per million characters, roughly half ElevenLabs’ rate and a multiple of budget options.
Production features target agent use: inline tags for non-verbal sounds like laughter, voice cloning from about 10 seconds of audio, custom pronunciation dictionaries with IPA support, and dependable reading of order numbers and confirmation codes. Hinglish demos point at Indian call-center deployments.
Three months after Sonic-3.5, the update arrives as enterprises weigh standardizing voice infrastructure on a single vendor.