← Search

Srikanth Ronanki

10 accepted papers

2025

Compress, Gather, and Recompute: REFORMing Long-Context Processing in Transformers

NeurIPS 2025poster

As large language models increasingly gain popularity in real-world applications, processing extremely long contexts, often exceeding the model’s pre-trained context limits, has emerged as a critical challenge. While existing approaches to efficient long-context processing show promise, recurrent co…

Cited by 0SourceScholar
2025

Context Length Alone Hurts LLM Performance Despite Perfect Retrieval

EMNLP 2025

Large language models (LLMs) often fail to scale their performance on long-context tasks performance in line with the context lengths they support. This gap is commonly attributed to retrieval failures—the models’ inability to identify information in the long inputs that is relevant to the task they

Cited by 0SourcePDFScholar
2025

LAWCAT: Efficient Distillation from Quadratic to Linear Attention with Convolution across Tokens for Long Context Modeling

EMNLP 2025

Although transformer architectures have achieved state-of-the-art performance across diverse domains, their quadratic computational complexity with respect to sequence length remains a significant bottleneck, particularly for latency-sensitive long-context applications. While recent linear-complexit

Cited by 0SourcePDFScholar
2025

Speech Retrieval-Augmented Generation without Automatic Speech Recognition

ICASSP 2025accepted

One common approach for question answering over speech data is to first transcribe speech using automatic speech recognition (ASR) and then employ text-based retrieval-augmented generation (RAG) on the transcriptions. While this cascaded pipeline has proven effective in many practical settings, ASR…

Cited by 0SourceScholar
2025

Zero-resource Speech Translation and Recognition with LLMs

ICASSP 2025accepted

Despite recent advancements in speech processing, zero-resource speech translation (ST) and automatic speech recognition (ASR) remain challenging problems. In this work, we propose to leverage a multilingual Large Language Model (LLM) to perform ST and ASR in languages for which the model has never…

Cited by 0SourceScholar
2024

SpeechGuard: Exploring the Adversarial Robustness of Multi-modal Large Language Models

ACL 2024findings

Integrated Speech and Large Language Models (SLMs) that can follow speech instructions and generate relevant text responses have gained popularity lately. However, the safety and robustness of these models remains largely unclear. In this work, we investigate the potential vulnerabilities of such in…

2023

Dynamic Chunk Convolution for Unified Streaming and Non-Streaming Conformer ASR

ICASSP 2023accepted

Recently, there has been an increasing interest in unifying streaming and non-streaming speech recognition models to reduce development, training and deployment cost. The best-known approaches rely on either window-based or dynamic chunk-based attention strategy and causal convolutions to minimize t…

Cited by 0SourceScholar
2021

Transformer-Transducers for Code-Switched Speech Recognition

ICASSP 2021accepted

We live in a world where 60% of the population can speak two or more languages fluently. Members of these communities constantly switch between languages when having a conversation. As automatic speech recognition (ASR) systems are being deployed to the real-world, there is a need for practical syst…

Cited by 0SourceScholar
2019

Effect of Data Reduction on Sequence-to-sequence Neural TTS

ICASSP 2019accepted

Recent speech synthesis systems based on sampling from autoregressive neural network models can generate speech almost indistinguishable from human recordings. However, these models require large amounts of data. This paper shows that the lack of data from one speaker can be compensated with data fr…

Cited by 0SourceScholar
2016

Robust TTS duration modelling using DNNS

ICASSP 2016accepted

Accurate modelling and prediction of speech-sound durations is an important component in generating more natural synthetic speech. Deep neural networks (DNNs) offer a powerful modelling paradigm, and large, found corpora of natural and expressive speech are easy to acquire for training them. Unfortu…

Cited by 0SourceScholar