← Search

Ashutosh Pandey

19 accepted papers

2026

ARRAYDPS-REFINE: GENERATIVE REFINEMENT OF DISCRIMINATIVE MULTI-CHANNEL SPEECH ENHANCEMENT

ICASSP 2026poster

Multi-channel speech enhancement aims to recover clean speech from noisy multi-channel recordings. Most deep learning methods employ discriminative training, which can lead to non-linear distortions from regression-based objectives, especially under challenging environmental noise conditions. Inspir…

Cited by 0SourcePDFScholar
2025

Advancing Active Speaker Detection for Egocentric Videos

ICASSP 2025accepted

This paper presents an improved approach to multimodal active speaker detection in egocentric videos, specifically designed to be robust against the rapid movements and motion blur commonly found in such videos. We propose two key techniques to improve the model’s resilience: (i) spatially fixing th…

Cited by 0SourceScholar
2025

Data-driven Processing using Parametric Neural Network for Improved Bluetooth Channel Sounding Distance Estimation

ICASSP 2025accepted

Accurate device-to-device distance estimation is crucial for Internet-of-things (IoT) applications. Traditional methods, such as RSSI-based ranging and Time-of-Flight narrowband systems, exhibit limitations. Bluetooth Low Energy (BLE)-based phase ranging, aka Channel Sounding is a preferred technolo…

Cited by 0SourceScholar
2025

Modulating State Space Model with SlowFast Framework for Compute-Efficient Ultra Low-Latency Speech Enhancement

ICASSP 2025accepted

Deep learning-based speech enhancement (SE) methods often face significant computational challenges when needing to meet low-latency requirements because of the increased number of frames to be processed. This paper introduces the SlowFast framework which aims to reduce computation costs specificall…

Cited by 0SourceScholar
2025

Reexamining the Efficacy of MetricGAN for Speech Enhancement

ICASSP 2025accepted

MetricGAN, a notable generative approach, provides an effective framework to train speech enhancement models to produce high metric scores. However, we identify two key limitations of current MetricGAN-family models, i.e. neglecting certain mainstream metrics during evaluation and conducting evaluat…

Cited by 0SourceScholar
2025

Robust Frame-level Speaker Localization in Reverberant and Noisy Environments by Exploiting Phase Difference Losses

ICASSP 2025accepted

This paper investigates robust speaker localization at the frame level on the basis of complex spectral mapping, which is capable of learning both the magnitude and phase of the target signal. Unlike prevailing deep learning methods for speaker localization, we perform MIMO (multi-input multi-output…

Cited by 0SourceScholar
2025

WiSenseNet: A Unified Foundation Model for Diverse Wi-Fi Sensing Tasks Using Channel State Information

ICASSP 2025accepted

Wi-Fi sensing utilizing Channel State Information (CSI) has emerged as a promising non-invasive technique for environmental perception, but current approaches are hindered by task-specific architectures, limited generalization, and data inefficiency, impeding its versatility across diverse applicati…

Cited by 0SourceScholar
2024

Decoupled Spatial and Temporal Processing for Resource Efficient Multichannel Speech Enhancement

ICASSP 2024accepted

We present a novel model designed for resource-efficient multichannel speech enhancement in the time domain, with a focus on low latency, lightweight, and low computational requirements. The proposed model incorporates explicit spatial and temporal processing within deep neural network (DNN) layers.…

Cited by 0SourceScholar
2024

Leveraging Sound Localization to Improve Continuous Speaker Separation

ICASSP 2024accepted

Continuous speaker separation aims to separate overlapping speakers in real-world environments like meetings, but it often falls short in isolating speech segments of a single speaker. This leads to split signals that adversely affect downstream applications such as automatic speech recognition and…

Cited by 0SourceScholar
2024

On the Importance of Neural Wiener Filter for Resource Efficient Multichannel Speech Enhancement

ICASSP 2024accepted

We introduce a time-domain framework for efficient multichannel speech enhancement, emphasizing low latency and computational efficiency. This framework incorporates two compact deep neural networks (DNNs) surrounding a multichannel neural Wiener filter (NWF). The first DNN enhances the speech signa…

Cited by 0SourceScholar
2024

WIFIACT: Enhancing Human Sensing Through Environment Robust Preprocessing And Bayesian Self-Supervised Learning

ICASSP 2024accepted

Wi-Fi Sensing is emerging as a transformative paradigm in the realm of smart environments, enabling the ubiquitous detection of human presence and the identification of activities within indoor spaces. This paper presents WiFiAct, which leverages a 20 MHz 1 transmit 1 receive (1T1R) Wi-Fi monitor to…

Cited by 0SourceScholar
2022

Multichannel Speech Enhancement Without Beamforming

ICASSP 2022accepted

Deep neural networks are often coupled with traditional spatial filters, such as MVDR beamformers for effectively exploiting spatial information. Even though single-stage end-to-end supervised models can obtain impressive enhancement, combining them with a traditional beamformer and a DNN-based post…

Cited by 0SourceScholar
2022

TPARN: Triple-Path Attentive Recurrent Network for Time-Domain Multichannel Speech Enhancement

ICASSP 2022accepted

In this work, we propose a new model called triple-path attentive recurrent network (TPARN) for multichannel speech enhancement in the time domain. TPARN extends a single-channel dual-path network to a multichannel network by adding a third path along the spatial dimension. First, TPARN processes sp…

Cited by 0SourceScholar
2020

Densely Connected Neural Network with Dilated Convolutions for Real-Time Speech Enhancement in The Time Domain

ICASSP 2020accepted

In this work, we propose a fully convolutional neural network for real-time speech enhancement in the time domain. The proposed network is an encoder-decoder based architecture with skip connections. The layers in the encoder and the decoder are followed by densely connected blocks comprising of dil…

Cited by 0SourceScholar
2019

TCNN: Temporal Convolutional Neural Network for Real-time Speech Enhancement in the Time Domain

ICASSP 2019accepted

This work proposes a fully convolutional neural network (CNN) for real-time speech enhancement in the time domain. The proposed CNN is an encoder-decoder based architecture with an additional temporal convolutional module (TCM) inserted between the encoder and the decoder. We call this architecture…

Cited by 0SourceScholar
2015

A novel Time-Delay-of-Arrival estimation technique for multi-microphone audio processing

ICASSP 2015accepted

Multi-microphone speech enhancement requires knowledge of relative Time Delay of Arrival (TDOA) of the desired acoustic source at microphones. This paper presents a novel TDOA estimation method, Steered Null Error PHAse Transform (SNE-PHAT), which exploits null-steering to improve estimation robustn…

Cited by 0SourceScholar