← Search

Sourav Bhattacharya

7 accepted papers

2026

HierarchicalPrune: Position-Aware Compression for Large-Scale Diffusion Models

AAAI 2026technical

State-of-the-art text-to-image diffusion models (DMs) achieve remarkable quality, yet their massive parameter scale (8-11B) poses significant challenges for inferences on resource-constrained devices. In this paper, we present HierarchicalPrune, a novel compression framework grounded in a key observ

Cited by 0SourcePDFScholar
2025

EDiT: Efficient Diffusion Transformers with Linear Compressed Attention

ICCV 2025poster

Diffusion Transformers (DiTs) have emerged as a leading architecture for text-to-image synthesis, producing high-quality and photorealistic images. However, the quadratic scaling properties of the attention in DiTs hinder image generation with higher resolution or devices with limited resources. Thi…

Cited by 0SourcePDFScholar
2025

Linear Time Complexity Conformers with SummaryMixing for Streaming Speech Recognition

ICASSP 2025accepted

Automatic speech recognition (ASR) with an encoder equipped with self-attention, whether streaming or non-streaming, takes quadratic time in the length of the speech utterance. This slows down training and decoding, increase the cost, and limits the deployment of the ASR in constrained devices. Summ…

Cited by 0SourceScholar
2025

Upcycling Text-to-Image Diffusion Models for Multi-Task Capabilities

ICML 2025poster

Text-to-image synthesis has witnessed remarkable advancements in recent years. Many attempts have been made to adopt text-to-image models to support multiple tasks. However, existing approaches typically require resource-intensive re-training or additional parameters to accommodate for the new tasks…

Cited by 0SourcePDFScholar
2024

MobileQuant: Mobile-friendly Quantization for On-device Language Models

EMNLP 2024finding

Large language models (LLMs) have revolutionized language processing, delivering outstanding results across multiple applications. However, deploying LLMs on edge devices poses several challenges with respect to memory, energy, and compute costs, limiting their widespread use in devices such as mobi…

2022

Conditioning Sequence-to-sequence Networks with Learned Activations

ICLR 2022poster

Conditional neural networks play an important role in a number of sequence-to-sequence modeling tasks, including personalized sound enhancement (PSE), speaker dependent automatic speech recognition (ASR), and generative modeling such as text-to-speech synthesis. In conditional neural networks, the o…

Cited by 13SourcePDFScholar
2021

NAS-Bench-ASR: Reproducible Neural Architecture Search for Speech Recognition

ICLR 2021poster

Powered by innovations in novel architecture design, noise tolerance techniques and increasing model capacity, Automatic Speech Recognition (ASR) has made giant strides in reducing word-error-rate over the past decade. ASR models are often trained with tens of thousand hours of high quality speech d…

Cited by 86SourcePDFScholar