← Search

Qianhua He

12 accepted papers

2026

Measuring Audio's Impact on Correctness: Audio-Contribution-Aware Post-Training of Large Audio Language Models

ICLR 2026poster

Large Audio Language Models (LALMs) represent an important frontier in multimodal AI, addressing diverse audio tasks. Recently, post-training of LALMs has received increasing attention due to significant performance improvements over foundation models. While single-stage post-training such as reinfo…

Cited by 0SourcecodeScholar
2025

An Efficient Sample Utilization Method for Deep Learning Based on Class Uncertainty

ICASSP 2025accepted

Deep learning has achieved success across many domains when sufficient training samples are available. However, the commonly used mini-batch stochastic gradient descent (SGD) training paradigm treats each sample equally, resulting in massive computational waste on samples that are easily identifiabl…

Cited by 0SourceScholar
2025

Cross-Domain Few-Shot Open-Set Keyword Spotting Using Keyword Adaptation and Prototype Reprojection

ICASSP 2025accepted

Personalized keyword spotting (KWS) with few enrollment utterances remains an important problem over years. KWS remains a challenging task due to the following factors, including the scarcity of enrollment samples, speech variation in the open-set scenarios, and distributional gap between source and…

Cited by 0SourceScholar
2025

Deep Enhancement Spotting Network for Low-complexity Keyword Spotting in Noisy Environments

ICASSP 2025accepted

Keyword Spotting (KWS) is crucial for hands-free voice-activated systems, requiring a balance between accuracy and complexity, especially in noisy environments. While Speech Enhancement (SE) can improve KWS accuracy, existing methods often lack the ability to effectively utilize the rich features pr…

Cited by 0SourceScholar
2023

Clean Sample Guided Self-Knowledge Distillation for Image Classification

ICASSP 2023accepted

For two-stage knowledge distillation, the combination with Data Augmentation (DA) is straightforward and effective. Yet, for online Self-knowledge Distillation (SD), DA is not always beneficial because of the absence of a trustworthy teacher model. To address this issue, this paper proposes an SD me…

Cited by 0SourceScholar
2021

A Stage Match for Query-by-Example Spoken Term Detection Based On Structure Information of Query

ICASSP 2021accepted

The state-of-the-art of query-by-example spoken term detection (QbE-STD) strategies are usually based on segmental dynamic time warping (S-DTW). However, the sliding window in S-DTW may separate signal of a word into different segments and produce many illegal candidates required to be compared with…

Cited by 0SourceScholar
2021

Domestic Activities Clustering From Audio Recordings Using Convolutional Capsule Autoencoder Network

ICASSP 2021accepted

Recent efforts have been made on domestic activities classification from audio recordings, especially the works submitted to the challenge of DCASE (Detection and Classification of Acoustic Scenes and Events) since 2018. In contrast, few studies were done on domestic activities clustering, which is…

Cited by 0SourceScholar
2017

Mobile phone clustering from acquired speech recordings using deep Gaussian supervector and spectral clustering

ICASSP 2017accepted

Acquisition device clustering from speech recordings is a new and critical problem in the field of speech forensic, which aims at merging speech recordings acquired by the same device into one cluster without both pre-knowing prior information of the processed data and pre-training classifier. We pr…

Cited by 0SourceScholar
2016

Source cell phone matching from speech recordings by sparse representation and KISS metric

ICASSP 2016accepted

Source recording device matching from two speech recordings is a new and important problem of digital media forensics. It aims to answer the question that whether or not two speech recordings are recorded by the same recording device. In this study we propose a source cell phone matching scheme. The…

Cited by 0SourceScholar
2015

Acoustic feature extraction by tensor-based sparse representation for sound effects classification

ICASSP 2015accepted

This paper describes a method to extract time-frequency (TF) audio features by tensor-based sparse approximation for sound effects classification. In the proposed method, the observed data is encoded as a higher-order tensor and discriminative features are extracted in spectrotemporal domain. Firstl…

Cited by 0SourceScholar