← Search

Aditya Gourav

6 accepted papers

2025

Align-SLM: Textless Spoken Language Models with Reinforcement Learning from AI Feedback

ACL 2025long

While textless Spoken Language Models (SLMs) have shown potential in end-to-end speech-to-speech modeling, they still lag behind text-based Large Language Models (LLMs) in terms of semantic coherence and relevance. This work introduces the Align-SLM framework, which leverages preference optimization…

2025

Speech Recognition Rescoring with Large Speech-Text Foundation Models

ICASSP 2025accepted

Large language models (LLM) have demonstrated the ability to understand human language by leveraging large amount of text data. Automatic speech recognition (ASR) systems are often limited by available transcribed speech data and benefit from a second pass rescoring using LLM. Recently multi-modal l…

Cited by 0SourceScholar
2024

Multi-Modal Retrieval For Large Language Model Based Speech Recognition

ACL 2024findings

Retrieval is a widely adopted approach for improving language models leveraging external information. As the field moves towards multi-modal large language models, it is important to extend the pure text based methods to incorporate other modalities in retrieval as well for applications across the w…

2021

Domain-Aware Neural Language Models for Speech Recognition

ICASSP 2021accepted

As voice assistants become more ubiquitous, they are increasingly expected to support and perform well on a wide variety of use-cases across different domains. We present a domain-aware rescoring framework suitable for achieving domain-adaptation during second-pass rescoring in production settings.…

Cited by 0SourceScholar
2021

Personalization Strategies for End-to-End Speech Recognition Systems

ICASSP 2021accepted

The recognition of personalized content, such as contact names, remains a challenging problem for end-to-end speech recognition systems. In this work, we demonstrate how first- and second-pass rescoring strategies can be leveraged together to improve the recognition of such words. Following previous…

Cited by 0SourceScholar