← Search

Jagadeesh Balam

14 accepted papers

2025

Anticipating Future with Large Language Model for Simultaneous Machine Translation

NAACL 2025long

Simultaneous machine translation (SMT) takes streaming input utterances and incrementally produces target text. Existing SMT methods only use the partial utterance that has already arrived at the input and the generated hypothesis. Motivated by human interpreters’ technique to forecast future words…

Cited by 0SourcePDFScholar
2025

Chain-of-Thought Prompting for Speech Translation

ICASSP 2025accepted

Large language models (LLMs) have demonstrated remarkable advancements in language understanding and generation. Building on the success of text-based LLMs, recent research has adapted these models to use speech embeddings for prompting, resulting in Speech-LLM models that exhibit strong performance…

Cited by 0SourceScholar
2025

Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data

ICASSP 2025accepted

Recent end-to-end speech language models (SLMs) have expanded upon the capabilities of large language models (LLMs) by incorporating pre-trained speech models. However, these SLMs often undergo extensive speech instruction-tuning to bridge the gap between speech and text modalities. This requires si…

Cited by 0SourceScholar
2025

EMMeTT: Efficient Multimodal Machine Translation Training

ICASSP 2025accepted

A rising interest in the modality extension of foundation language models warrants discussion on the most effective, and efficient, multimodal training approach. This work focuses on neural machine translation (NMT) and proposes a joint multimodal training regime of Speech-LLM to include automatic s…

Cited by 0SourceScholar
2025

META-CAT: Speaker-Informed Speech Embeddings via Meta Information Concatenation for Multi-talker ASR

ICASSP 2025accepted

We propose a novel end-to-end multi-talker automatic speech recognition (ASR) framework that enables both multi-speaker (MS) ASR and target-speaker (TS) ASR. Our proposed model is trained in a fully end-to-end manner, incorporating speaker supervision from a pre-trained speaker diarization module. W…

Cited by 0SourceScholar
2025

NEST: Self-supervised Fast Conformer as All-purpose Seasoning to Speech Processing Tasks

ICASSP 2025accepted

Self-supervised learning (SSL) has been proved to benefit a wide range of speech processing tasks, such as speech recognition/translation, speaker verification and diarization, etc. However, most of current speech SSL approaches are computationally expensive. In this paper, we introduce a simplified…

Cited by 0SourceScholar
2025

Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

ICML 2025poster

Sortformer is an encoder-based speaker diarization model designed for supervising speaker tagging in speech-to-text models. Instead of relying solely on permutation invariant loss (PIL), Sortformer introduces Sort Loss to resolve the permutation problem, either independently or in tandem with PIL. I…

Cited by 0SourcePDFScholar
2025

VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning

NAACL 2025long

Recent studies have augmented large language models (LLMs) with speech capabilities, leading to the development of speech language models (SpeechLMs). Earlier SpeechLMs focused on single-turn speech-based question answering (QA), where user input comprised a speech context and a text question. More…

2024

Discrete Audio Representation as an Alternative to Mel-Spectrograms for Speaker and Speech Recognition

ICASSP 2024accepted

Discrete audio representation, aka audio tokenization, has seen renewed interest driven by its potential to facilitate the application of text language modeling approaches in audio domain. To this end, various compression and representation-learning based tokenization schemes have been proposed. How…

Cited by 0SourceScholar
2024

Enhancing Speaker Diarization with Large Language Models: A Contextual Beam Search Approach

ICASSP 2024accepted

Large language models (LLMs) have shown great promise for capturing contextual information in natural language processing tasks. We propose a novel approach to speaker diarization that incorporates the prowess of LLMs to exploit contextual cues in human dialogues. Our method builds upon an acoustic-…

Cited by 0SourceScholar
2024

Investigating End-to-End ASR Architectures for Long Form Audio Transcription

ICASSP 2024accepted

This paper presents an overview and evaluation of some of the end-to-end ASR models on long-form audio. We study three categories of Automatic Speech Recognition(ASR) models based on their core architecture: (1) convolutional, (2) convolutional with squeeze-and-excitation, and (3) convolutional mode…

Cited by 0SourceScholar
2024

Multilingual Audio-Visual Speech Recognition with Hybrid CTC/RNN-T Fast Conformer

ICASSP 2024accepted

Humans are adept at leveraging visual cues from lip movements for recognizing speech in adverse listening conditions. Audio-Visual Speech Recognition (AVSR) models follow similar approach to achieve robust speech recognition in noisy conditions. In this work, we present a multilingual AVSR model inc…

Cited by 0SourceScholar
2024

SALM: Speech-Augmented Language Model with in-Context Learning for Speech Recognition and Translation

ICASSP 2024accepted

We present a novel Speech Augmented Language Model (SALM) with multitask and in-context learning capabilities. SALM comprises a frozen text LLM, a audio encoder, a modality adapter module, and LoRA layers to accommodate speech input and associated task instructions. The unified SALM not only achieve…

Cited by 0SourceScholar
2024

Stateful Conformer with Cache-Based Inference for Streaming Automatic Speech Recognition

ICASSP 2024accepted

In this paper, we propose an efficient and accurate streaming speech recognition model based on the FastConformer architecture. We adapted the FastConformer architecture for streaming applications through: (1) constraining both the look-ahead and past contexts in the encoder, and (2) introducing an…

Cited by 0SourceScholar