← Search

Xuanjun Chen

5 accepted papers

2026

How Does Instrumental Music Help SingFake Detection?

ICASSP 2026poster

Although many models exist to detect singing voice deepfakes (SingFake), how these models operate, particularly with instrumental accompaniment, is unclear. We investigate how instrumental music affects SingFake detection from two perspectives. To investigate the behavioral effect, we test different…

Cited by 0SourcePDFScholar
2025

Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks

ICLR 2025poster

Multimodal foundation models, such as Gemini and ChatGPT, have revolutionized human-machine interactions by seamlessly integrating various forms of data. Developing a universal spoken language model that comprehends a wide range of natural language instructions is critical for bridging communication…

2025

Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech Enhancement

ICASSP 2025accepted

In multichannel speech enhancement, effectively capturing spatial and spectral information across different microphones is crucial for noise reduction. Traditional methods, such as CNN or LSTM, attempt to model the temporal dynamics of full-band and sub-band spectral and spatial features. However, t…

Cited by 0SourceScholar
2024

Codec-SUPERB: An In-Depth Analysis of Sound Codec Models

ACL 2024findings

The sound codec’s dual roles in minimizing data transmission latency and serving as tokenizers underscore its critical importance.Recent years have witnessed significant developments in codec models.The ideal sound codec should preserve content, paralinguistics, speakers, and audio information.Howev…

2024

Multimodal Transformer Distillation for Audio-Visual Synchronization

ICASSP 2024accepted

Audio-visual synchronization aims to determine whether the mouth movements and speech in the video are synchronized. VocaLiST reaches state-of-the-art performance by incorporating multimodal Transformers to model audio-visual interact information. However, it requires high computing resources, makin…

Cited by 0SourceScholar