← Search

Arnav Kundu

8 accepted papers

2026

EpiCache: Episodic KV Cache Management for Long-Term Conversation on Resource-Constrained Environments

ICML 2026poster

Modern large language models (LLMs) extend context lengths to millions of tokens, enabling coherent, personalized responses grounded in long conversational history. However, the Key-Value (KV) cache grows linearly with the extended dialogue history, causing the model’s memory footprint to quickly ex…

Cited by 0SourceScholar
2026

MemoryLLM: Plug-n-Play Interpretable Feed-Forward Memory for Transformers

ICML 2026poster

Understanding how transformer components operate in LLMs is important, as it is at the core of recent technological advances in artificial intelligence. In this work, we revisit the challenges associated with interpretability of feed-forward modules (FFNs) and propose MemoryLLM, which aims to decoup…

Cited by 0SourceScholar
2025

An Efficient and Streaming Audio Visual Active Speaker Detection System

ICASSP 2025accepted

This paper delves into the challenging task of active speaker detection (asd), where the system needs to determine in real-time whether a person is speaking or not in a series of video frames. While previous works have made significant strides in improving network architectures and learning effectiv…

Cited by 0SourceScholar
2025

SPD: Sync-Point Drop for Efficient Tensor Parallelism of Large Language Models

ICML 2025poster

With the rapid expansion in the scale of large language models (LLMs), enabling efficient distributed inference across multiple computing units has become increasingly critical. However, communication overheads from popular distributed inference techniques such as Tensor Parallelism pose a significa…

Cited by 0SourcePDFScholar
2024

Streaming Anchor Loss: Augmenting Supervision with Temporal Significance

ICASSP 2024accepted

Streaming neural network models for fast frame-wise responses to various speech and sensory signals are widely adopted on resource-constrained platforms. Hence, increasing the learning capacity of such streaming models (i.e., by adding more parameters) to improve the predictive power may not be viab…

Cited by 0SourceScholar
2023

HEiMDaL: Highly Efficient Method for Detection and Localization of Wake-Words

ICASSP 2023accepted

Streaming keyword spotting is a widely used solution for activating voice assistants. Methods based on Deep Neural Networks with Hidden Markov Model (DNN-HMM) have proven to be efficient and widely adopted in this space, primarily because of the ability to detect and identify the start and end of th…

Cited by 0SourceScholar
2023

I See What You Hear: A Vision-Inspired Method to Localize Words

ICASSP 2023accepted

This paper explores the possibility of using visual object detection techniques for word localization in speech data. Object detection has been thoroughly studied in the contemporary literature for visual data. Noting that an audio can be interpreted as a 1-dimensional image, object localization tec…

Cited by 0SourceScholar
2021

Optimize What Matters: Training DNN-Hmm Keyword Spotting Model Using End Metric

ICASSP 2021accepted

Deep Neural Network–Hidden Markov Model (DNN-HMM) based methods have been successfully used for many always-on keyword spotting algorithms that detect a wake word to trigger a device. The DNN predicts the state probabilities of a given speech frame, while HMM decoder combines the DNN predictions of…

Cited by 0SourceScholar