Interfaze Ships diffusion-gemma-asr-small, an Open-Source Diffusion ASR Model Transcribing Six Languages via DiffusionGemma’s Parallel Denoising Decoder
Interfaze, a younger YC’s startup, has open-sourced a brand new speech recognition mannequin. It is known as diffusion-gemma-asr-small. The mannequin transcribes audio by way of a diffusion decoder, not an autoregressive one. It is described as the primary multilingual audio diffusion ASR mannequin. One adapter handles six languages. The analysis workforce educated solely about 42M…
