Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas
Cartesia has launched Sonic-3.6, the latest model of its real-time text-to-speech mannequin. It arrives roughly three months after Sonic-3.5. The new change is naturalness, and this one is independently checkable. Sonic 3.6 now holds #1 on each Artificial Analysis speech leaderboards — 1,283 Elo on the Provider Voice board and 1,123 on the Controlled Voice board. The second consequence issues extra. That board clones each mannequin onto the identical eight reference voices, which isolates the synthesis engine from the voice catalog. Sonic-3.6 leads it, with Sonic-3.5 second and ElevenLabs Eleven v3 third. The mannequin runs on state space models somewhat than transformers, and Cartesia states sub-90ms time-to-first-audio. It is obtainable in beta.
Is it deployable?
YES, it’s obtainable in beta and as a hosted API. Not as self-hosted weights.
Sonic is a closed, business mannequin. There are not any open weights and no Hugging Face repo. You lease it.
- Company stage: Solo builders and startups (Free/Pro $5 tiers), scaleups operating contact facilities (Startup $49 / Scale $299), and controlled enterprises needing DPAs, BAAs, and SSO.
- Industries: Financial providers, healthcare, retail and e-commerce, logistics, recruiting, SaaS help, client companion apps, media localization
- Applications: Inbound help brokers, outbound qualification calls, IVR substitute, appointment reminders, sales-training simulators, audio localization, in-product voice UI
The Architecture
Sonic runs on state area fashions somewhat than transformers. Cartesia’s launch page frames the same old tradeoffs — velocity versus naturalness, accuracy versus price — as architectural, not inevitable.
The sensible output is time-to-first-audio. Cartesia states sub-90ms TTS latency, and 100ms transcript latency for its Ink-2 speech-to-text mannequin. Both are vendor-stated mannequin latency, not measured end-to-end spherical journeys.
Interactive explainer
Features that matter in manufacturing
Sonic exposes controls constructed for agent transcripts somewhat than narration:
- Inline expression tags. Non-verbal expressions like
[laughter]go straight within the transcript. - Instant voice cloning from about 10 seconds of audio.
- Custom pronunciation dictionaries, together with IPA overrides comparable to
<<s|ə|ˈ|p|i|n|ə>>for subpoena. - Speed, quantity, and emotion parameters uncovered by way of the API and integrations just like the LiveKit Agents plugin.
- Native alphanumerics. Order numbers, cellphone numbers, and affirmation codes learn accurately with out preprocessing.
Cartesia’s launch demos present English with pure pauses and filler phrases, plus Hinglish code-switching between Hindi and English.
Pricing actuality
Artificial Analysis normalizes Sonic 3.6 at $49.00 per 1M characters. That is half of ElevenLabs Eleven v3 at $100.00, and effectively above Speechify Simba 3.2 at $10.00 for a 1,240 Elo.
Cartesia sells credit, not characters. Scale at $299 per 30 days contains roughly 10,667 TTS minutes and 15 concurrent requests. Line voice agents invoice individually at $0.06 per minute.
Key Takeaways
- Sonic-3.6 is #1 on each Artificial Analysis speech arenas — 1,283 Elo Provider Voice, 1,123 Controlled Voice.
- Winning the Controlled board means the engine improved, not simply the voice catalog.
- It is beta on Cartesia’s API solely; docs nonetheless listing Sonic 3.5 as steady, and companions carry 3.5.
- Deployable as a hosted API, not self-hosted weights; business use begins on the $5 Pro tier.
- Latency claims (sub-90ms TTFA) are vendor-stated mannequin latency, so benchmark your individual spherical journey.
Check out the Project Page-Cartesia Sonic, Cartesia launch page, Cartesia pricing, Cartesia docs, Artificial Analysis Speech Arena and @cartesia on X. All figures verified August 18, 2026.. Also, be at liberty to observe us on Twitter and don’t overlook to hitch our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to accomplice with us for selling your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar and many others.? Connect with us
The publish Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas appeared first on MarkTechPost.
