LLMs Learn to Clean Up Noisy Speech Recognition with Less Data
Researchers show that large language models can be efficiently fine-tuned to correct speech recognition errors caused by background noise, achieving up to a 53.9% improvement in word error rate correction.
A joint research team from NVIDIA Research Taiwan, the MIT-IBM Watson AI Lab, Nanyang Technological University, Georgia Institute of Technology, the University of Aberdeen, and Tsinghua University has demonstrated that large language models can be efficiently fine-tuned to correct speech recognition errors caused by background noise. Presented at ICLR 2024 on May 7, the work shows that LLMs can perform generative error correction for automatic speech recognition under noisy conditions, achieving up to a 53.9% improvement in word error rate correction compared to baseline generative error correction approaches. The method requires only limited training data, a departure from the data-intensive pipelines that dominate current noise-robust ASR research.
The work addresses a persistent bottleneck in real-world voice technology. Automatic speech recognition systems have become highly accurate in clean, quiet conditions, but performance degrades sharply in environments such as factories, call centers, train stations, and busy streets. Conventional approaches to noise robustness typically demand massive, diverse audio datasets and specialized acoustic models. The new research shows that an LLM can learn to denoise transcripts within its own linguistic space, using a compact representation of noise extracted from the ASR system itself. This could lower the cost and complexity of deploying reliable voice interfaces across industries and geographies.
Closing the Cross-Modality Gap
The central technical challenge is what the researchers call the “cross-modality gap.” Audio-derived noise embeddings, which mathematically represent the acoustic conditions of a recording, are structurally different from the token embeddings that LLMs process. When raw audio embeddings are fed directly into an LLM during fine-tuning, they can disrupt the model’s internal representations and degrade performance. The team’s solution is a new type of representation: a language-space noise embedding.
This embedding is extracted from the N-best hypotheses list generated by the ASR system. The N-best list contains multiple candidate transcriptions ranked by likelihood. The pattern of disagreement among these candidates encodes useful information about where and how noise corrupted the recognition. The researchers use knowledge distillation to transfer real noise information from audio embeddings into this language-space embedding. The result is a vector that is compatible with the LLM’s token space, allowing the model to condition its corrections on the actual acoustic conditions without being disrupted by foreign feature types.
The approach builds on the HyPoradise benchmark for generative error correction, extending it specifically to noisy scenarios. In the paper, the authors report that the method achieves the 53.9% WER correction improvement against baseline generative error correction approaches, a figure that underscores the practical value of aligning noise representations with the LLM’s native modality. By operating entirely in language space, the model avoids the need for direct audio processing during fine-tuning, which simplifies the training pipeline and reduces computational overhead.
Efficiency as a Design Principle
A striking aspect of the research is its emphasis on data efficiency. The team achieved the reported gains using limited training data, a deliberate departure from the data-hungry paradigm that dominates much of modern ASR research. This matters for organizations that cannot collect millions of hours of labeled audio across every possible noise condition. Instead of training a new acoustic model from scratch, the method fine-tunes a pre-trained LLM to act as a post-processor, correcting errors in the ASR output.
The authors include YuChen Hu, Chen Chen, Ruizhe Li, Chao Zhang, and EnSiong Chng from Nanyang Technological University; Huck Yang from Georgia Institute of Technology and NVIDIA Research; and Pin-Yu Chen from the MIT-IBM Watson AI Lab. The collaboration also involved the University of Aberdeen and Tsinghua University. This multi-institutional effort reflects the growing recognition that robust speech interfaces require expertise spanning natural language processing, acoustic modeling, and efficient machine learning.
By treating noise as a property that can be represented in language space, the model learns to denoise without explicit acoustic supervision during fine-tuning. The LLM’s pre-trained linguistic knowledge provides a strong prior for what plausible text looks like, enabling it to reject implausible corrections even when the audio signal is heavily degraded. This synergy between large-scale language pretraining and compact noise conditioning is the core insight of the work. The limited data requirement also suggests that the method could be applied in settings where collecting extensive noisy speech corpora is impractical, such as low-resource languages or specialized industrial domains.
Implications for Global Voice Interfaces
For enterprises and developers building voice assistants, transcription platforms, or voice-controlled industrial systems, the findings suggest a more accessible path to robustness. Current noise-robust ASR pipelines often require specialized acoustic front-ends, domain-specific data collection, and careful tuning for each deployment environment. The new method offers an alternative: keep the existing ASR system largely intact, and add an LLM-based correction layer that adapts to noise through a learned embedding.
The potential applications extend across sectors. In call centers, where background chatter and varying microphone quality degrade transcripts, an LLM-based correction module could improve downstream analytics and compliance monitoring. In manufacturing or logistics, voice commands issued near machinery could be interpreted more reliably, reducing error rates in hands-free operations. In public services and healthcare, transcription of noisy field recordings could become more accurate without requiring expensive re-recording or manual review.
The method’s data efficiency also has implications for languages and dialects with limited training resources. Collecting large noisy speech corpora for every language is impractical. If an LLM can learn noise correction from a modest amount of data in one language and transfer that capability across its multilingual pretraining, the barrier to robust ASR in low-resource languages could drop significantly. The paper does not claim full cross-lingual transfer, but the architectural principles are compatible with multilingual LLMs.
Reactions and Research Context
The ICLR 2024 presentation places this work within a broader trend of using LLMs as post-hoc correctors for ASR systems. Earlier generative error correction methods showed promise in clean conditions but faltered when noise was introduced. The new research directly addresses that weakness, and the reported improvement over baselines is substantial enough to attract attention from both academic and industrial groups working on speech interfaces.
Pin-Yu Chen, a researcher at the MIT-IBM Watson AI Lab and one of the paper’s authors, has previously published on robustness and efficiency in machine learning. The collaboration with NVIDIA Research Taiwan reflects the company’s ongoing investment in speech and language technologies, particularly in contexts where models must operate under real-world constraints. The work was also highlighted in connection with IBM’s presence at AAAI 2024, where related themes of efficient and robust AI were discussed.
It is important to note that the results are based on benchmark evaluations under specific noise conditions. Real-world deployment will introduce additional variables, including reverberation, overlapping speech, and code-switching. The paper does not claim to solve all of these, and the authors acknowledge that generative error correction is one component of a larger ASR pipeline. Nonetheless, the demonstrated efficiency and the magnitude of the improvement suggest that language-space noise embeddings could become a standard tool in the speech recognition toolkit.
Looking ahead, the research opens several directions. One is the extension of language-space noise embeddings to streaming ASR, where latency constraints require corrections to be generated incrementally. Another is the combination of this method with parameter-efficient fine-tuning techniques such as LoRA or adapters, which could further reduce the computational cost of adapting LLMs to new noise environments. The team’s focus on limited training data aligns with a growing industry emphasis on sustainable, cost-effective AI development.
As voice becomes a primary interface for AI systems, the ability to understand speech in noisy, unpredictable environments will determine how widely these systems are adopted. The NVIDIA and MIT-IBM collaboration demonstrates that large language models, already transforming text generation and understanding, can also serve as efficient, adaptive denoisers for the spoken word. By translating acoustic noise into the language of the model itself, the researchers have charted a practical route toward voice AI that works not just in quiet labs, but in the noisy world where people actually live and work.
Sources
- Large Language Models are Efficient Learners of Noise-Robust Speech Recognition | NVIDIA Research Taiwan
- IBM at AAAI 2024 - Vancouver, BC, Canada
Written by an AI editorial process from the sources above. Errors may occur.
Newsletter
Get the AI news that matters
One short brief with the day's most important AI stories — written for professionals.
We send a confirmation link. No spam. Unsubscribe anytime.
Read next
AI Accelerates Drug Discovery from Concept to Clinic
Artificial intelligence is slashing development timelines and costs in pharmaceutical R&D, with AI-designed drugs now entering clinical trials in record time.
27 Sep 2026
Google Moves Gemini Team Under DeepMind Leadership
Google integrates its consumer AI app team into DeepMind to accelerate generative AI development and streamline research-to-product pipelines.
25 Sep 2026
AI in Drug Discovery: From Target ID to Clinical Trials
Artificial intelligence is accelerating drug discovery, but clinical validation remains the final frontier.
24 Sep 2026