← Search

Mingru Yang

4 accepted papers

2026

Measuring Audio's Impact on Correctness: Audio-Contribution-Aware Post-Training of Large Audio Language Models

ICLR 2026poster

Large Audio Language Models (LALMs) represent an important frontier in multimodal AI, addressing diverse audio tasks. Recently, post-training of LALMs has received increasing attention due to significant performance improvements over foundation models. While single-stage post-training such as reinfo…

Cited by 0SourcecodeScholar
2025

An Efficient Sample Utilization Method for Deep Learning Based on Class Uncertainty

ICASSP 2025accepted

Deep learning has achieved success across many domains when sufficient training samples are available. However, the commonly used mini-batch stochastic gradient descent (SGD) training paradigm treats each sample equally, resulting in massive computational waste on samples that are easily identifiabl…

Cited by 0SourceScholar
2025

Cross-Domain Few-Shot Open-Set Keyword Spotting Using Keyword Adaptation and Prototype Reprojection

ICASSP 2025accepted

Personalized keyword spotting (KWS) with few enrollment utterances remains an important problem over years. KWS remains a challenging task due to the following factors, including the scarcity of enrollment samples, speech variation in the open-set scenarios, and distributional gap between source and…

Cited by 0SourceScholar
2025

Deep Enhancement Spotting Network for Low-complexity Keyword Spotting in Noisy Environments

ICASSP 2025accepted

Keyword Spotting (KWS) is crucial for hands-free voice-activated systems, requiring a balance between accuracy and complexity, especially in noisy environments. While Speech Enhancement (SE) can improve KWS accuracy, existing methods often lack the ability to effectively utilize the rich features pr…

Cited by 0SourceScholar