← Search

Yanzhen Ren

11 accepted papers

2026

Profiling the Voice: Speaker-Specific Phoneme Fingerprinting for Speech Deepfake Detection

IJCAI 2026

The rapid advancement of generative AI has made audio deepfakes increasingly indistinguishable from authentic human vocals, posing significant threats to persons-of-interest (POI) such as public figures. Current detection systems primarily rely on generic, black-box models that fail to capture speak

Cited by 0Scholar
2026

When AVSR Meets Video Conferencing: Dataset, Degradation, and the Hidden Mechanism Behind Performance Collapse

CVPR 2026

Audio-Visual Speech Recognition (AVSR) has achieved remarkable progress in offline conditions, yet its robustness in real-world video conferencing (VC) remains largely unexplored. This paper presents the first systematic evaluation of state-of-the-art AVSR models across mainstream VC platforms, reve

Cited by 0SourceScholar
2025

Attention Weighting and Conditional Entropy-driven Quantization Loss for Neural Audio Codecs

ICASSP 2025accepted

Existing end-to-end neural codecs have made great progress in preserving audio quality. Despite their success, they still face challenges in achieving accurate and efficient quantization. Specifically, these codecs often overlook which features have a greater impact on perceptual audio quality durin…

Cited by 0SourceScholar
2025

FreqSense: Universal and Low-Latency Adversarial Example Detection for Speaker Recognition with Interpretability in Frequency Domain

ICASSP 2025accepted

Speaker recognition (SR) systems are particularly vulnerable to adversarial example (AE) attacks. To mitigate these attacks, AE detection systems are typically integrated into SR systems. To overcome the limitations of low detection accuracy, poor generalization, and high latency in existing schemes…

Cited by 0SourceScholar
2025

Improving Speech Enhancement by Cross- and Sub-band Processing with State Space Model

ICASSP 2025accepted

Recently, the state space model (SSM) represented by Mamba has shown remarkable performance in long-term sequence modeling tasks, including speech enhancement. However, due to substantial differences in sub-band features, applying the same SSM to all sub-bands limits its inference capability. Additi…

Cited by 0SourceScholar
2024

FCC-MF: Detecting Violence in Audio-Visual Context with Frame-Wise Cluster Contrast and Modality-Stage Flooding

ICASSP 2024accepted

This paper explores the detection of frame-wise instances of violence in both audio and visual modalities, where only clip-level labels are available. Previous works selected fixed value of frames for objective optimization to model frame-level features, and applied straightforward fusion strategy t…

Cited by 0SourceScholar
2024

Semantic Proximity Alignment: Towards Human Perception-Consistent Audio Tagging by Aligning with Label Text Description

ICASSP 2024accepted

Most audio tagging models are trained with one-hot labels as supervised information. However, one-hot labels treat all sound events equally, ignoring the semantic hierarchy and proximity relationships between sound events. In contrast, the event descriptions contains richer information, describing t…

Cited by 0SourceScholar
2024

VFD-Net: Vocoder Fingerprints Detection for Fake Audio

ICASSP 2024accepted

With the rapid development of audio deepfake technology, the credibility and authenticity of public opinion is facing a formidable challenge. Since vocoder is the key component of audio deepfake and leaves distinctive fingerprint features, we propose VFD-Net (Vocoder Fingerprints Detection Net), a n…

Cited by 0SourceScholar
2023

Attention Mixup: An Accurate Mixup Scheme Based On Interpretable Attention Mechanism for Multi-Label Audio Classification

ICASSP 2023accepted

Mixup proves to be an efficient data augmentation method on audio classification tasks. Original mixup scheme directly mixes the waveform of two random samples, which not only ignores the temporal distribution of the sound events but may also interfere with the original sound events in another sampl…

Cited by 0SourceScholar
2023

Learning From Single-Expert Annotated Labels for Automatic Sleep Staging

ICASSP 2023accepted

Existing automatic sleep staging algorithms rely on accurately labeled data. However, due to the subjectivity of sleep experts, accurate labels must be obtained through joint labeling by multiple experts, which results in high time and labor costs. In this work, we treat labels mislabeled by a singl…

Cited by 0SourceScholar