← Search

Han Yin

7 accepted papers

2026

Can Large Audio Language Models Understand Audio Well? Speech, Scene and Events Understanding Benchmark for LALMs

ICASSP 2026poster

Recently, Large Audio Language Models (LALMs) have progressed rapidly, demonstrating their strong efficacy in universal audio understanding through cross-modal integration. To evaluate LALMs' audio understanding performance, researchers have proposed different benchmarks. However, key aspects for re…

Cited by 0SourcePDFScholar
2026

Environmental Sound Deepfake Detection Challenge: An Overview

ICASSP 2026poster

Recent progress in audio generation models has made it possible to create highly realistic and immersive soundscapes, which are now widely used in film and virtual-reality-related applications. However, these audio generators also raise concerns about potential misuse, such as producing deceptive au…

Cited by 0SourcePDFScholar
2026

MOESCORE: MIXTURE-OF-EXPERTS-BASED TEXT-AUDIO RELEVANCE SCORE PREDICTION FOR TEXT-TO-AUDIO SYSTEM EVALUATION

ICASSP 2026poster

Recent advances in generative models have enabled modern Text-to-Audio (TTA) systems to synthesize audio with high perceptual quality. However, TTA systems often struggle to maintain semantic consistency with the input text, leading to mismatches in sound events, temporal tructures, or contextual re…

Cited by 0SourcePDFScholar
2025

Exploring Text-Queried Sound Event Detection with Audio Source Separation

ICASSP 2025accepted

In sound event detection (SED), overlapping sound events pose a significant challenge, as certain events can be easily masked by background noise or other events, resulting in poor detection performance. To address this issue, we propose the text-queried SED (TQ-SED) framework. Specifically, we firs…

Cited by 0SourceScholar
2024

Multimodal Clickbait Detection by De-confounding Biases Using Causal Representation Inference

EMNLP 2024main

This paper focuses on detecting clickbait posts on the Web. These posts often use eye-catching disinformation in mixed modalities to mislead users to click for profit. That affects the user experience and thus would be blocked by content provider. To escape detection, malicious creators use tricks t…

Cited by 0SourcePDFScholar
2023

3D Audio Signal Processing Systems for Speech Enhancement and Sound Localization and Detection

ICASSP 2023accepted

The L3DAS23 of ICASSP Signal Processing Grand Challenge encourages research on 3D audio signal processing, such as 3D speech enhancement (SE) and 3D sound localization and detection (SELD). In this paper, we propose a two-stage system based on DPRNN and UNet for the SE task and a Conformer-based sys…

Cited by 0SourceScholar