|

Best Open Speech Recognition (ASR) Models in 2026: WER, Languages, Latency, and License Compared

Open speech recognition stopped being a Whisper monoculture a while in the final twelve months. In March 2026 Cohere released Transcribe, a 2B Apache 2.0 mannequin that took the highest of the Hugging Face Open ASR Leaderboard at 5.42% common phrase error fee. Five weeks later IBM shipped Granite Speech 4.1 2B at 5.33%. Since then ARK-ASR-3B and MOSS-Transcribe-preview-2B have posted decrease numbers nonetheless.

The high of that leaderboard is now separated by lower than one WER level. That has a particular consequence for anybody selecting a mannequin: rank is not the deciding variable. License, language protection, streaming assist, and price per audio-hour are. This roundup compares the sphere on all 4.


First, an issue with the leaderboard quantity everybody quotes

The Open ASR Leaderboard common is just not a single mounted amount, and the fashions at the moment listed aspect by aspect weren’t all scored the identical method.

Cohere’s 5.42% is a median throughout eight English check units, together with TED-LIUM. That checks out: AMI 8.13, Earnings-22 10.86, GigaSpeech 9.34, LibriSpeech clear 1.25, LibriSpeech different 2.37, SPGISpeech 3.08, TED-LIUM 2.49, VoxPopuli 5.87 averages to precisely 5.42.

ARK-ASR-3B’s 5.04% is a median throughout seven units. TED-LIUM is absent. The MOSS-Transcribe-preview-2B card states this explicitly: TED-LIUM is just not at the moment a part of the leaderboard run and is subsequently excluded.

TED-LIUM is likely one of the simpler units in the suite, so dropping it raises the typical. Recompute Cohere’s revealed per-dataset scores over the identical seven units ARK reviews and Cohere lands at 5.84, not 5.42. Do the identical to Granite Speech 4.1 2B and it strikes from 5.33 to five.65. On a like-for-like foundation ARK’s lead is bigger than the headline numbers suggest, not smaller — however the level is that you simply can’t subtract one revealed determine from one other and get a significant reply.

Two additional caveats belong on the identical web page:

Some scores are brazenly leaderboard-fitted: The MOSS-Transcribe-preview-2B card states the mannequin was fine-tuned with reinforcement studying on the Open ASR Leaderboard coaching splits. That is disclosed, which is greater than most, but it surely means the rating measures the benchmark reasonably than the potential.

Private-track information reorders the board: Appen contributed held-back evaluation sets masking Australian, Canadian, Indian, and American accents in scripted and conversational circumstances. When these personal units are toggled on, zoom/scribe_v1 strikes from #4 to #1 and the public-leaderboard chief drops a place. Models tuned for clear learn speech degrade disproportionately on spontaneous conversational audio.

Use the leaderboard to construct a shortlist. Do not use it to choose a winner.

The accuracy tier

Cohere Transcribe (2B, Apache 2.0, 14 languages) is the mannequin that truly shipped into manufacturing. It has been downloaded over 620,000 instances in the previous month and has runtime assist throughout transformers, vLLM, mlx-audio for Apple Silicon, a Rust port, and a WebGPU construct. It is a Conformer encoder with a light-weight Transformer decoder, educated from scratch. Cohere additionally ran human choice analysis, the place educated annotators scored transcripts for which means preservation, hallucination, and named entities: a 61% common win fee, 78% in opposition to IBM Granite 4.0 1B Speech and 64% in opposition to Whisper large-v3.

The limitations part of its mannequin card is unusually sincere and needs to be learn earlier than committing. There isn’t any computerized language detection, no timestamps, and no diarization, and the mannequin is raring to transcribe silence, so Cohere recommends prepending a VAD or noise gate. The repo can be gated behind a contact-information settlement regardless of the Apache 2.0 license.

Granite Speech 4.1 2B (2B, Apache 2.0) is the higher choose should you want functionality reasonably than a decrease quantity. Six languages for ASR plus bidirectional speech translation, keyword-list biasing for names and jargon, punctuation and truecasing together with German noun capitalization. Trained on 174,000 hours. RTFx 231.29. IBM additionally ships two siblings: -plus provides speaker-attributed ASR and word-level timestamps, and -nar is mentioned under.

Canary-Qwen-2.5B (2.5B, CC-BY-4.0, English) pairs a QuickConformer encoder with a Qwen3-1.7B decoder and runs in two modes — pure transcription, or LLM mode the place the decoder summarizes and solutions questions concerning the transcript. 5.63% WER at RTFx 418. Note that AMI was oversampled to roughly 15% of coaching information, which biases output towards verbatim disfluency-preserving transcripts. That is a function for authorized work and a nuisance for assembly notes.

Qwen3-ASR-1.7B (Apache 2.0) covers 52 languages and dialects — 30 languages plus 22 Chinese dialects — at 5.76%. It ships with a full inference toolkit and a separate forced-alignment mannequin for timestamps in 11 languages. For something touching Mandarin or Chinese regional speech that is the apparent place to begin.

The throughput tier

Accuracy throughout the highest of the sphere now varies by about one WER level. Throughput varies by greater than an order of magnitude, which suggests throughput often decides the bill.

Parakeet TDT 0.6B v3 (0.6B, CC-BY-4.0) is the throughput chief amongst multilingual open fashions at RTFx 3332.74 throughout 25 European languages with computerized language ID, dealing with as much as 24 minutes in a single go on an A100 80GB. It prices 6.32% WER — roughly one level greater than Granite 4.1 2B for roughly fourteen instances the audio per GPU-second.

Granite Speech 4.1 2B-NAR is the extra fascinating engineering outcome. It is non-autoregressive: it edits a CTC speculation in a single ahead go utilizing a bidirectional LLM, reaching RTFx ~1820 on one H100 at batch measurement 128. It provides up Japanese, speech translation, and key phrase biasing to get there.

Qwen3-ASR-0.6B retains all 52 languages and reaches 2000× throughput at concurrency 128.

The streaming tier

A batch WER appears to be the improper check for a streaming mannequin, and the leaderboard scores them anyway. Voxtral Realtime sits at 7.68% and Kyutai STT 2.6B at 6.40% — each under Whisper — and neither quantity tells you something helpful about their supposed use.

Voxtral Mini 4B Realtime 2602 (Apache 2.0, 13 languages) is a 3.4B language mannequin together with a 970M causal audio encoder educated from scratch, with sliding-window consideration on each halves for successfully unbounded streaming. Transcription delay is configurable in 80ms steps from 80ms to 1200ms, plus a standalone 2400ms possibility; Mistral recommends 480ms because the candy spot and reviews that at that setting it matches main offline open fashions. It runs on a single 16GB GPU and had day-0 vLLM Realtime API support.

Kyutai STT (CC-BY-4.0) comes in two shapes: a ~1B English/French mannequin with a 0.5s delay and a built-in semantic voice exercise detector, and a 2.6B English-only mannequin with a 2.5s delay. For voice brokers, the semantic VAD issues greater than the transcription delay — it predicts when the speaker has really completed, which is what governs perceived turn-taking latency. An H100 serves 400 concurrent streams in actual time.

The protection tier

Meta’s Omnilingual ASR (Apache 2.0, corpus CC-BY) is just not competing on WER and shouldn’t be evaluated as if it have been. It covers 1,600+ languages natively and extends to five,400+ by way of zero-shot in-context studying, constructed on a wav2vec 2.0 encoder scaled to 7B and pre-trained on about 4.3M hours. The 7B LLM-ASR variant achieves character error fee under 10% on 78% of supported languages, together with 500+ by no means beforehand served by any ASR system. Encoder sizes run from 300M to 7B. Meta additionally launched the Omnilingual ASR Corpus masking 350+ underserved languages.

Whisper large-v3 (1.55B, MIT, 99 languages) has been overtaken on accuracy by roughly ten open fashions, and it stays the proper default for a big class of initiatives. MIT is the least encumbered license in the sphere. The runtime ecosystem — whisper.cpp, faster-whisper, WhisperX — has no equal among the many newer releases. If your requirement is “some language, some {hardware}, no license lawyer,” it’s nonetheless the reply.

The analysis edge

diffusion-gemma-asr-small from YC startup Interfaze is essentially the most architecturally inventive launch of the 12 months. It produces transcripts by parallel diffusion denoising over a 256-token canvas in 8 to 16 steps, so decoding price doesn’t develop with transcript size. Only ~42M parameters have been educated — 0.16% of the weights — on high of a frozen 26B DiffusionGemma and a frozen whisper-small encoder. It reaches 6.6% WER on LibriSpeech test-clean at roughly 11 to 17× realtime. It additionally reaches 15.7% on FLEURS English and 29.6% CER on FLEURS Mandarin, so deal with the LibriSpeech determine because the ceiling. Nobody ought to deploy this. Everybody engaged on ASR ought to learn it.

MOSS-Transcribe-Diarize 0.9B (Apache 2.0, 50+ languages) solves the issue most roundups ignore: it emits speaker labels, phrase timestamps, and transcript in one technology as a substitute of chaining ASR to a separate diarization stack. 128k context, roughly 90 minutes of audio in one go, RTF ~0.017 on an RTX 4090, with hotword biasing.

The license cut up no one reads till deployment

This is the part that truly blocks transport, and the sphere divides cleanly:

Apache 2.0: Cohere Transcribe, Granite Speech 4.1 (all three variants), Qwen3-ASR (each sizes), Voxtral Mini Realtime, Omnilingual ASR, ARK-ASR, MOSS-Transcribe. No attribution obligation, business use unrestricted. Note that Cohere’s repo is gated behind a contact-information settlement though the license itself is Apache 2.0.

MIT: Whisper large-v3. The most permissive possibility in the sphere.

CC-BY-4.0: Canary-Qwen-2.5B, Parakeet TDT 0.6B v3, Kyutai STT. Commercially usable, however attribution is required. For an embedded product or a white-labeled API, that may be a actual compliance obligation, and it’s the single commonest cause groups find yourself transport a mannequin that isn’t essentially the most correct one they examined.

Meta’s Omnilingual ASR splits the 2: Apache 2.0 for fashions, CC-BY for the corpus.

How to truly select

Run this order, not simply the leaderboard order:

  1. License: If attribution is a blocker, CC-BY-4.0 removes Parakeet, Canary-Qwen, and Kyutai earlier than you benchmark something.
  2. Language protection: Cohere’s 14 languages, Granite’s 6, and Canary’s English-only are laborious limits, not tender ones. Cohere moreover has no language auto-detection, so you need to know the language in advance.
  3. Streaming or batch: This is architectural. No quantity of tuning turns an offline encoder-decoder right into a low-latency streaming mannequin.
  4. Then measure WER by yourself audio: The unfold between the highest ten fashions on the general public benchmark is beneath one level. The unfold in your accented, noisy, domain-specific audio shall be a number of instances that, and it is not going to rank the fashions in the identical order.
  5. Then compute price per audio-hour by yourself GPUs: RTFx figures are measured at massive batch sizes on datacenter {hardware} and don’t switch.


The helpful abstract of 2026 is just not that any single mannequin received. It is {that a} 2B open-weight mannequin on a permissive license now beats what closed APIs have been charging for eighteen months in the past, and that the remaining resolution is a procurement query reasonably than a analysis one.


Sources:

Leaderboard positions are dwell and change often. All figures verified in opposition to major sources on 23 July 2026.

The put up Best Open Speech Recognition (ASR) Models in 2026: WER, Languages, Latency, and License Compared appeared first on MarkTechPost.

Similar Posts