**arXiv ID:** 2604.14001 **Authors:** Davyd Naveriani, Albert Zeyer, Ralf Schlüter, Hermann Ney **Published:** 2026-04-15T15:46:15Z **Abstract:** Diffusion language models have recently emerged as a leading alternative to standard language models, due to their ability for bidirectional attention and parallel text generation. In this work, we explore variants for their use in speech recognition. Specifically, we introduce a comprehensive guide to incorporating masked diffusion language models ...
Scanned 9/11/2026
Install to Claude Code
npx -y skills add hiyenwong/ai_collection --skill diffusion-language-models-for-speech-recognition --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Diffusion Language Models For Speech Recognition?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/hiyenwong-diffusion-language-models-for-speech-recognition)More formats (shields.io, HTML) on the badges page.
# Diffusion Language Models for Speech Recognition
**arXiv ID:** 2604.14001
**Authors:** Davyd Naveriani, Albert Zeyer, Ralf Schlüter, Hermann Ney
**Published:** 2026-04-15T15:46:15Z
**Abstract:**
Diffusion language models have recently emerged as a leading alternative to standard language models, due to their ability for bidirectional attention and parallel text generation. In this work, we explore variants for their use in speech recognition. Specifically, we introduce a comprehensive guide to incorporating masked diffusion language models (MDLM) and uniform-state diffusion models (USDMs) for rescoring ASR hypotheses. Additionally, we design a new joint-decoding method that combines CTC and USDM by integrating the framewise probability distributions derived from CTC with the labelwise probability distributions computed by USDM at each decoding step, thereby generating new candidates that combine strong language knowledge from USDM and acoustic information from CTC. Our findings reveal that USDM, as well as MDLM, can significantly improve the accuracy of recognized text. We publish all our code and recipes.
## Skill Description
This skill is generated from the arXiv paper: Diffusion Language Models for Speech Recognition (2604.14001).
## How to Use
[To be filled in by the user or by future automation]
## References
- [arXiv:2604.14001](http://arxiv.org/abs/2604.14001v2)
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!