← Search

Gongping Huang

33 accepted papers

2026

CoCoEmo: Composable and Controllable Human-Like Emotional TTS via Activation Steering

ICML 2026poster

Emotional expression in human speech is nuanced and compositional, often involving multiple, sometimes conflicting, affective cues that may diverge from linguistic content. In contrast, most expressive text-to-speech (TTS) systems enforce a single utterance-level emotion, collapsing affective divers…

Cited by 0SourceScholar
2026

DTT-BSR: GAN-BASED DTTNET WITH ROPE TRANSFORMER ENHANCEMENT FOR MUSIC SOURCE RESTORATION

ICASSP 2026poster

Music source restoration (MSR) aims to recover unprocessed stems from mixed and mastered recordings. The challenge lies in both separating overlapping sources and reconstructing signals degraded by production effects such as compression and reverberation. We therefore propose DTT-BSR, a hybrid gener…

Cited by 0SourcePDFScholar
2026

FORWARD CONVOLUTIVE PREDICTION FOR FRAME ONLINE MONAURAL SPEECH DEREVERBERATION BASED ON KRONECKER PRODUCT DECOMPOSITION

ICASSP 2026poster

Dereverberation has long been a crucial research topic in speech processing, aiming to alleviate the adverse effects of reverberation in voice communication and speech interaction systems. Among existing approaches, forward convolutional prediction (FCP) has recently attracted attention. It typicall…

Cited by 0SourcePDFScholar
2026

MEANFLOW-ACCELERATED MULTIMODAL VIDEO-TO-AUDIO SYNTHESIS VIA ONE-STEP GENERATION

ICASSP 2026poster

A key challenge in synthesizing audios from silent videos is the inherent trade-off between synthesis quality and inference efficiency in existing methods. For instance, flow matching based models rely on modeling instantaneous velocity, inherently require an iterative sampling process, leading to s…

Cited by 0SourcePDFScholar
2026

ONLINE NEURAL FUSION OF DISTORTIONLESS DIFFERENTIAL BEAMFORMERS FOR ROBUST SPEECH ENHANCEMENT

ICASSP 2026poster

Fixed beamforming is widely used in practice since it does not depend on the estimation of noise statistics and provides relatively stable performance. However, a single beamformer cannot adapt to varying acoustic conditions, which limits its interference suppression capability. To address this, ada…

Cited by 0SourcePDFScholar
2026

ROBUST ONLINE OVERDETERMINED INDEPENDENT VECTOR ANALYSIS BASED ON BILINEAR DECOMPOSITION

ICASSP 2026oral

Online blind source separation is essential for both speech communication and human-machine interaction. Among existing approaches, overdetermined independent vector analysis (OverIVA) delivers strong performance by exploiting the statistical independence of source signals and the orthogonality betw…

Cited by 0SourcePDFScholar
2026

Semantic Audio-Visual Navigation in Continuous Environments

CVPR 2026

Audio-visual navigation enables embodied agents to navigate toward sound-emitting targets by leveraging both auditory and visual cues. However, most existing approaches rely on precomputed room impulse responses (RIRs) for binaural audio rendering, restricting agents to discrete grid positions and l

Cited by 0SourcecodeScholar
2025

Advances in Microphone Array Processing and Multichannel Speech Enhancement

ICASSP 2025accepted

This paper reviews pioneering works in microphone array processing and multichannel speech enhancement, highlighting historical achievements, technological evolution, commercialization aspects, and key challenges. It provides valuable insights into the progression and future direction of these areas…

Cited by 0SourceScholar
2025

DOA Estimation Based on Enhanced SRP-MVDR Using Kronecker Product Decomposition for Large Rectangular Microphone Arrays

ICASSP 2025accepted

Direction-of-arrival (DOA) estimation is a key process in microphone array systems. The steered response power-based minimum variance distortionless response (SRP-MVDR) method performs very well in challenging acoustic environments but suffers from exponential complexity as the number of microphones…

Cited by 0SourceScholar
2025

Data-Driven White Noise Gain Constrained Robust Superdirective Beamformer for Speech Enhancement

ICASSP 2025accepted

Superdirective beamformers are highly effective at suppressing directional interference and diffuse noise, but their practical use is often constrained by the problem of white noise amplification. Robust superdirective beamforming methods typically address this by imposing a constraint on the white…

Cited by 0SourceScholar
2025

Design and Optimization of Superdirective Beamforming and Post-Filtering for Speech Enhancement

ICASSP 2025accepted

Superdirective beamformers, used with small microphone arrays, are highly attractive due to their high directivity and frequency-invariant beampatterns, making them well-suited for processing broadband acoustic and speech signals. However, these beamformers are very sensitive to array imperfections…

Cited by 0SourceScholar
2025

Design of Robust Differential Beamformers with Microphone Arrays of Arbitrary Planar Geometry

ICASSP 2025accepted

Differential microphone arrays (DMAs) have garnered significant attention in recent research and development due to their high directivity and frequency-invariant beampatterns. However, DMAs frequently encounter substantial white noise amplification, which limits their practical applications. This p…

Cited by 0SourceScholar
2025

LMFCA-Net: A Lightweight Model for Multi-Channel Speech Enhancement with Efficient Narrow-Band and Cross-Band Attention

ICASSP 2025accepted

Deep learning based end-to-end multi-channel speech enhancement methods have achieved impressive performance by leveraging sub-band, cross-band, and spatial information. However, these methods often demand substantial computational resources, limiting their practicality on terminal devices. This pap…

Cited by 0SourceScholar
2025

Microphone Array Beamforming for Speech Enhancement Based on Dynamic Mode Decomposition

ICASSP 2025accepted

Microphone array beamforming is widely used to extract desired speech signals from noisy environments. While most research in this area focuses on utilizing spatial information, less attention is given to the intrinsic physical mechanisms underlying microphone array observations. This paper aims to…

Cited by 4SourceScholar
2025

On the Design of Low-Rank Differential Beamformers with Nonuniform Linear Microphone Arrays

ICASSP 2025accepted

Kronecker product beamforming is an effective technique for designing beamformers with nonuniform linear arrays (NULAs). However, current techniques are restricted to NULAs with specific configurations, where the steering vector of the array is represented as a Kronecker product of steering vectors…

Cited by 0SourceScholar
2024

Beamforming Through Online Convex Combination of Differential Beamformers

ICASSP 2024accepted

Thanks to their high directivity, compact size, and reliable performance, differential microphone arrays (DMAs) have attracted great interest from both industry and academia as they have demonstrated great potential to be used in a wide range of applications for high-fidelity speech acquisition. Nev…

Cited by 0SourceScholar
2024

Differential Beamforming with Null Constraints for Spherical Microphone Arrays

ICASSP 2024accepted

Differential microphone arrays (DMAs) can measure both the acoustic pressure field and the differential acoustic pressure fields, which gives them great advantages in a wide range of applications for acoustic and speech signal acquisition. The core component of DMAs is the so-called differential bea…

Cited by 0SourceScholar
2024

On the Design of Planar Differential Microphone Arrays with Specified Beamwidth or Sidelobe Level

ICASSP 2024accepted

This paper investigates the problem of designing differential beam-formers with planar microphone arrays to achieve not only the desired target directivity pattern but also control the beamwidth (BW) or sidelobe level (SLL). We first discuss the target directivity patterns and express the Dolph-Cheb…

Cited by 0SourceScholar
2023

Spatially Informed Independent vector analysis for Source Extraction based on the convolutive Transfer Function Model

ICASSP 2023accepted

Spatial information can help improve source separation performance. Numerous spatially informed source extraction methods based on the independent vector analysis (IVA) have been developed, which can achieve reasonably good performance in non- or weakly reverberant environments. However, the perform…

Cited by 0SourceScholar
2023

Switching Kronecker Product Linear Filtering for Multispeaker Adaptive Speech Dereverberation

ICASSP 2023accepted

Dereverberation, a process to mitigate or eliminate the reverberation effect, plays an important role in hands-free speech communication and human-machine interfaces. Tremendous efforts have been devoted to this problem and various methods have been developed over the last three decades. Those metho…

Cited by 0SourceScholar
2021

Combined Differential Beamforming With Uniform Linear Microphone Arrays

ICASSP 2021accepted

While differential beamformers have been widely used in voice communication and human-machine speech interface systems to enhance speech signals of interest, how to design such beamformers that on the one hand can achieve the highest possible directivity factor (DF) and on the other hand are able to…

Cited by 0SourceScholar
2021

On the Design of Square Differential Microphone Arrays with a Multistage Structure

ICASSP 2021accepted

This paper studies the problem of designing square differential microphone arrays (SDMAs). It presents a multistage approach, which first divides an SDMA composed of M <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> microphones into (M − 1) <sup…

Cited by 0SourceScholar
2021

Robust Steerable Differential Beamformers with Null Constraints for Concentric Circular Microphone Arrays

ICASSP 2021accepted

Differential beamformers with concentric circular microphone arrays (CCMAs) are desirable for use in various applications since they can form frequency-invariant spatial responses, have better beam steering flexibility than linear arrays, and suffer less with beampattern irregularity and white noise…

Cited by 0SourceScholar
2020

An Improved Solution to the Frequency-Invariant Beamforming with Concentric Circular Microphone Arrays

ICASSP 2020accepted

Frequency-invariant beamforming with circular microphone arrays (CMAs) has drawn a significant amount of attention for its steering flexibility and high directivity. However, frequency-invariant beam-forming with CMAs often suffers from the so-called null problem, which is caused by the zeros of the…

Cited by 0SourceScholar
2020

Robust and steerable kronecker product differential beamforming With rectangular microphone arrays

ICASSP 2020accepted

Differential microphone arrays (DMAs), a class of welldesigned small-size arrays combined with differential beamforming, are very useful for processing broadband acoustic, audio, and speech signals in a wide range of applications. However, most efforts in the literature so far have been devoted to l…

Cited by 36SourceScholar
2019

Design of Optimal Linear Differential Microphone Arrays Based Array Geometry Optimization

ICASSP 2019accepted

This paper presents a method to design optimal linear differential microphone arrays (DMAs) by optimizing the array geometry. By constraining the DMA beamformer to achieve a given target value of the directivity factor (DF) with a specified target frequency-invariant beampattern while achieving also…

Cited by 0SourceScholar
2019

On the Design of Flexible Kronecker Product Beamformers with Linear Microphone Arrays

ICASSP 2019accepted

This paper proposes a method for the design of flexible Kronecker product beamformers based on the decomposition of the steering vector of a physical array as a Kronecker product of steering vectors of two smaller virtual arrays. With this decomposition, the global beamforming filter is designed by…

Cited by 0SourceScholar
2019

Properties and Limits of the Minimum-norm Differential Beamformers with Circular Microphone Arrays

ICASSP 2019accepted

Small aperture circular microphone arrays (CMAs) have been widely used in many applications such as teleconferencing, smartspeakers, and robotics. A critical component of such arrays is the differential beamformer, which can achieve relatively high spatial gains with the same beampatterns at most fr…

Cited by 0SourceScholar
2018

On the Design of Robust Steerable Frequency-Invariant Beampatterns with Concentric Circular Microphone Arrays

ICASSP 2018accepted

This paper studies the problem of frequency-invariant beamforming with concentric circular microphone arrays (CCMAs). We develop a beamforming algorithm based on an optimal approximation of the beamformer's beampattern with the Jacobi-Anger expansion. In comparison with the existing frequency-invari…

Cited by 25SourceScholar
2017

Study of the frequency-domain multichannel noise reduction problem with the householder transformation

ICASSP 2017accepted

This paper presents an approach to the multichannel noise reduction problem. It first transforms the multichannel noisy speech signals into the frequency domain. A Householder transformation is then constructed, which converts the multichannel coefficients in each frequency bin into two components:…

Cited by 3SourceScholar
2016

Subspace superdirective beamformers based on joint diagonalization

ICASSP 2016accepted

Although they have been intensively studied and used in many applications due to their high directivity factor (DF), superdirective beamformers are sensitive to sensor noise and mismatch between sensors. This paper studies the problem of superdirective beamforming combined with the joint diagonaliza…

Cited by 0SourceScholar
2015

Investigation of a parametric gain approach to single-channel speech enhancement

ICASSP 2015accepted

This paper investigates a parametric gain approach to single-channel noise reduction in the frequency domain. In comparison with the traditional parametric Wiener gain, the major novelty of this presented approach is that the parametric gain is formulated to estimate the noise by using the mean-squa…

Cited by 0SourceScholar
2015

Optimal single-channel noise reduction filtering matrices from the pearson correlation coefficient perspective

ICASSP 2015accepted

This paper studies the problem of single-channel noise reduction in the time domain, where an estimate of a vector of the desired clean speech is achieved by filtering a frame of the noisy signal with a rectangular filtering matrix. The core issue with this problem formulation is then the estimation…

Cited by 0SourceScholar