← Search

Qiquan Zhang

12 accepted papers

2026

ADAPTIVE PER-CHANNEL ENERGY NORMALIZATION FRONT-END FOR ROBUST AUDIO SIGNAL PROCESSING

ICASSP 2026poster

In audio signal processing, learnable front-ends have shown strong performance across diverse tasks by optimizing task-specific representation. However, their parameters remain fixed once trained, lacking flexibility during inference and limiting robustness under dynamic complex acoustic environment…

Cited by 0SourcePDFScholar
2026

ANALYTIC INCREMENTAL LEARNING FOR SOUND SOURCE LOCALIZATION WITH IMBALANCE RECTIFICATION

ICASSP 2026poster

Sound source localization (SSL) demonstrates remarkable results in controlled settings but struggles in real-world deployment due to dual imbalance challenges: intra-task imbalance arising from long-tailed direction-of-arrival (DoA) distributions, and inter-task imbalance induced by cross-task skews…

Cited by 0SourcePDFScholar
2026

BENCHMARKING GASLIGHTING ATTACKS AGAINST SPEECH LARGE LANGUAGE MODELS

ICASSP 2026poster

As Speech Large Language Models (Speech LLMs) become increasingly integrated into voice-based applications, ensuring their robustness against manipulative or adversarial input becomes critical. Although prior work has studied adversarial attacks in text-based LLMs and vision-language models, the uni…

Cited by 0SourcePDFScholar
2025

SpeechT-RAG: Reliable Depression Detection in LLMs with Retrieval-Augmented Generation Using Speech Timing Information

ACL 2025finding

Large Language Models (LLMs) have been increasingly adopted for health-related tasks, yet their performance in depression detection remains limited when relying solely on text input. While Retrieval-Augmented Generation (RAG) typically enhances LLM capabilities, our experiments indicate that traditi…

Cited by 0SourcePDFScholar
2025

Time-Graph Frequency Representation with Singular Value Decomposition for Neural Speech Enhancement

ICASSP 2025accepted

Time-frequency (T-F) domain methods for monaural speech enhancement have benefited from the success of deep learning. Recently, focus has been put on designing two-stream network models to predict amplitude mask and phase separately, or, coupling the amplitude and phase into Cartesian coordinates an…

Cited by 0SourceScholar
2024

An Empirical Study on the Impact of Positional Encoding in Transformer-Based Monaural Speech Enhancement

ICASSP 2024accepted

Transformer architecture has enabled recent progress in speech enhancement. Since Transformers are position-agostic, positional encoding is the de facto standard component used to enable Transformers to distinguish the order of elements in a sequence. However, it remains unclear how positional encod…

Cited by 0SourceScholar
2024

GLMB 3D Speaker Tracking with Video-Assisted Multi-Channel Audio Optimization Functions

ICASSP 2024accepted

Speaker tracking plays a significant role in numerous real-world human robot interaction (HRI) applications. In recent years, there has been a growing interest in utilizing multi-sensory information, such as complementary audio and visual signals, to address the challenges of speaker tracking. Despi…

Cited by 0SourceScholar
2024

Speaking in Wavelet Domain: A Simple and Efficient Approach to Speed up Speech Diffusion Model

EMNLP 2024main

Recently, Denoising Diffusion Probabilistic Models (DDPMs) have attained leading performances across a diverse range of generative tasks. However, in the field of speech synthesis, although DDPMs exhibit impressive performance, their prolonged training duration and substantial inference costs hinder…

2024

When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection

EMNLP 2024main

Depression is a critical concern in global mental health, prompting extensive research into AI-based detection methods. Among various AI technologies, Large Language Models (LLMs) stand out for their versatility in healthcare applications. However, the application of LLMs in the identification and a…

Cited by 11SourcePDFScholar
2023

Ripple Sparse Self-Attention for Monaural Speech Enhancement

ICASSP 2023accepted

The use of Transformer represents a recent success in speech enhancement. However, as its core component, self-attention suffers from quadratic complexity, which is computationally prohibited for long speech recordings. Moreover, it allows each time frame to attend to all time frames, neglecting the…

Cited by 10SourceScholar
2022

FineD-Eval: Fine-grained Automatic Dialogue-Level Evaluation

EMNLP 2022main

Recent model-based reference-free metrics for open-domain dialogue evaluation exhibit promising correlations with human judgment. However, they either perform turn-level evaluation or look at a single dialogue quality dimension. One would expect a good evaluation metric to assess multiple quality di…

2022

Time-Frequency Attention for Monaural Speech Enhancement

ICASSP 2022accepted

Most studies on speech enhancement generally don’t explicitly consider the energy distribution of speech in time-frequency (T-F) representation, which is important for accurate prediction of mask or spectra. In this paper, we present a simple yet effective T-F attention (TFA) module, where a…

Cited by 35SourceScholar