|

Gradium AI Releases New Default TTS Model: 81.0% Hard-Case Pass Rate at 216 ms Time-to-First-Audio

Voice brokers fail on precisely the components of a name that matter most: the order quantity, the callback digits, the e-mail deal with the caller has to jot down down. Gradium AI has launched a brand new text-to-speech mannequin and made it the default throughout its API and Studio. The firm experiences an 81.0% human-rated move charge on a 500-sentence hard-case set spanning 5 languages, forward of Cartesia Sonic 3.6 at 75.1% and ElevenLabs v3 Conversational at 65.4%. Time to first audio is 216 ms at P50 on Coval, 170 ms quicker than the mannequin it replaces.

Is it deployable?

Yes, at present, with no migration. Gradium switched the model on because the default throughout its API and Studio on August 31, 2026. Existing voices, together with customized clones, preserve working unchanged.


The accuracy quantity

Gradium constructed a 500-sentence analysis set and open-sourced it on Hugging Face underneath CC BY 4.0: 100 gadgets throughout 10 standards in 5 languages (EN, DE, FR, ES, PT). Seven atomic standards cowl spelling, acronyms, alphanumeric tokens, dates, common numbers, massive and floating numbers, and e mail. Three composite standards (Orders, IT Ticket, Claims) stack a number of of these into one real looking agent flip.

Scoring is human and strict. A sentence passes provided that an unbiased native-speaker rater hears each factor pronounced accurately and utterly; one dropped digit fails the sentence. Audio was loudness-normalized, order randomized, and raters capped at 40 comparisons with an enforced break.

Pooled throughout the ten standards and averaged over the 5 languages with equal weight: Gradium TTS 81.0%, Cartesia Sonic 3.6 75.1%, ElevenLabs v3 Conversational 65.4%, Fish Audio S2.1 Pro 49.5%, Inworld TTS 1.5 Max 46.5%. All generated in August 2026 with default settings.

The latency quantity

On Coval’s TTS benchmark, Gradium experiences a 216 ms P50 time to first audio, 170 ms quicker than the mannequin it replaces. The extra helpful determine is the unfold: a 30 ms p75-p25 interquartile vary throughout 480 runs, the tightest of the 5 fashions examined. Cartesia Sonic 3.6 sits at 454 ms median with a 165 ms unfold, 36% of its personal median, and callers expertise tail turns relatively than medians.

Gradium is just not the quickest mannequin on that chart. Inworld TTS 2 posts a 166 ms median; Fish Audio S2.1 Pro (291 ms) and ElevenLabs v3 Conversational (329 ms) path Gradium. The declare being made is about joint place: the bottom hard-case failure charge at sub-250 ms first audio, with little or no variance.

Getting began

Existing customers want do nothing. New groups set up the Python SDK, level at the WebSocket TTS endpoint, and reuse current voice IDs. Gradium is providing 1M credit for full hard-case failure experiences on its Discord.

Key Takeaways

  • New Gradium TTS mannequin is reside and default as of August 31, 2026; no migration wanted.
  • 81.0% human-rated move charge on 500 laborious sentences, forward of Cartesia, ElevenLabs, Fish Audio and Inworld.
  • 216 ms P50 time to first audio on Coval, with a 30 ms interquartile unfold over 480 runs.
  • Reads cellphone numbers, emails, IBANs and reference codes with no textual content normalization required.
  • Vendor-run benchmark, however the 500-sentence analysis set is open on Hugging Face underneath CC BY 4.0.


Check out the release post and the dataset. Also, be at liberty to observe us on Twitter and don’t neglect to hitch our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

Need to associate with us for selling your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar and so on.? Connect with us

The publish Gradium AI Releases New Default TTS Model: 81.0% Hard-Case Pass Rate at 216 ms Time-to-First-Audio appeared first on MarkTechPost.

Similar Posts