← Search

Zehua Zhang

11 accepted papers

2025

Hybrid Feature Global Attention Network for Noisy-reverberant Speech Enhancement

ICASSP 2025accepted

Deep neural network-based speech enhancement methods have become widespread, with one of its fundamental aspects being the effective extraction and application of features in the time-frequency domain. This paper proposes a hybrid feature global attention network (HFGANet) designed to efficiently ex…

Cited by 0SourceScholar
2025

PriorSinger: Singing Voice Synthesis Model with Prior Condition Cross Attention

ICASSP 2025accepted

The singing voice synthesis system is designed to generate realistic and expressive singing based on a given musical score. Generative Adversarial Networks (GANs) or diffusion models generate acoustic features, such as Mel-spectrograms, which are subsequently reconstructed into waveforms by a vocode…

Cited by 0SourceScholar
2024

Embracing Events and Frames with Hierarchical Feature Refinement Network for Object Detection

ECCV 2024poster

"In frame-based vision, object detection faces substantial performance degradation under challenging conditions due to the limited sensing capability of conventional cameras. Event cameras output sparse and asynchronous events, providing a potential solution to solve these problems. However, effecti…

2024

Hybrid Attention Time-Frequency Analysis Network for Single-Channel Speech Enhancement

ICASSP 2024accepted

The time-frequency domain remains central to the speech signal analysis. Enhancing the efficacy of neural network-based speech models demands a detailed multi-scale analysis of time-frequency features. This study presents the Hybrid Attention Time-Frequency Analysis Network (HATFANet), an innovative…

Cited by 0SourceScholar
2024

Lightweight Multi-Axial Transformer with Frequency Prompt for Single Channel Speech Enhancement

ICASSP 2024accepted

Time-frequency analysis in single-channel speech enhancement has received considerable attention. While Transformer-based architectures are gaining traction, their computational burden can be substantial, especially when dealing with longer speech samples. To address this, our research introduces th…

Cited by 0SourceScholar
2023

Half-Temporal and Half-Frequency Attention U2Net for Speech Signal Improvement

ICASSP 2023accepted

During communication, volume changes, noise, and reverberation can disturb speech signals, significantly affecting the quality and intelligibility of speech. In the context of the ICASSP 2023 Signal Processing Grand Challenge, the first Speech Signal Improvement Grand Challenge (SIG) is organized to…

Cited by 0SourceScholar
2023

Two-Stage UNet with Multi-Axis Gated Multilayer Perceptron for Monaural Noisy-Reverberant Speech Enhancement

ICASSP 2023accepted

In denoising and de-reverberation tasks, the dominant methods are complex spectral masking and complex spectral mapping. To combine advantages and improve speech enhancement performance, we propose a two-stage UNet (TSUNet) to estimate complex spectral masking and complex spectral mapping. We use a…

Cited by 0SourceScholar
2022

FB-MSTCN: A Full-Band Single-Channel Speech Enhancement Method Based on Multi-Scale Temporal Convolutional Network

ICASSP 2022accepted

In recent years, deep learning-based approaches have significantly improved the performance of single-channel speech enhancement. However, due to the limitation of training data and computational complexity, real-time enhancement of full-band (48 kHz) speech signals is still very challenging. Becaus…

Cited by 0SourceScholar
2020

Interaction Graphs for Object Importance Estimation in On-road Driving Videos

ICRA 2020poster

A vehicle driving along the road is surrounded by many objects, but only a small subset of them influence the driver's decisions and actions. Learning to estimate the importance of each object on the driver's real-time decision-making may help better understand human driving behavior and lead to mor…

Cited by 34SourceScholar
2020

Kalman Filtering Attention for User Behavior Modeling in CTR Prediction

NeurIPS 2020spotlight

Click-through rate (CTR) prediction is one of the fundamental tasks for e-commerce search engines. As search becomes more personalized, it is necessary to capture the user interest from rich behavior data. Existing user behavior modeling algorithms develop different attention mechanisms to emphasize…

2019

A Self Validation Network for Object-Level Human Attention Estimation

NeurIPS 2019poster

Due to the foveated nature of the human vision system, people can focus their visual attention on a small region of their visual field at a time, which usually contains only a single object. Estimating this object of attention in first-person (egocentric) videos is useful for many human-centered rea…