Show HN: The First Open Source Diffusion ASR Audio Model 15x Faster Than Whisper
By khurdula · 2026-07-21 · 2 points · 0 comments
https://arxiv.org/abs/2607.13013
We trained diffusion-gemma-asr, an open-source speech recognition model that is 15x faster than Whisper, based on DiffusionGemma and Whisper Small. Instead of generating text one token at a time like Whisper, in diffusion-gemma-asr we transcribe by denoising an entire transcript…
Open the full discussion on BetterNews