← Search

Haonan Cheng

8 accepted papers

2026

Detect All-Type Deepfake Audio: Wavelet Prompt Tuning for Enhanced Auditory Perception

AAAI 2026technical

The rapid advancement of audio generation technologies has escalated the risks of malicious deepfake audio across speech, sound, singing voice, and music, threatening multimedia security and trust. While existing countermeasures (CMs) perform well in single-type audio deepfake detection (ADD), their

Cited by 0SourcePDFScholar
2026

FAKE SPEECH WILD: DETECTING DEEPFAKE SPEECH ON SOCIAL MEDIA PLATFORM

ICASSP 2026poster

The rapid advancement of speech generation technology has led to the widespread proliferation of deepfake speech across social media platforms. While deepfake audio countermeasures (CMs) achieve promising results on public datasets, their performance degrades significantly in cross-domain scenarios.…

Cited by 0SourcePDFScholar
2025

Look Around Before Locating: Considering Content and Structure Information for Visual Grounding

AAAI 2025technical

As a long-term challenge and fundamental requirement in vision and language tasks, visual grounding aims to localize a target referred by a natural language query. The regional annotations form a superficial correlation between the subject of expression and some common visual entities, which hinder…

2024

An Efficient Temporary Deepfake Location Approach Based Embeddings for Partially Spoofed Audio Detection

ICASSP 2024accepted

Partially spoofed audio detection is a challenging task, lying in the need to accurately locate the authenticity of audio at the frame level. To address this issue, we propose a fine-grained partially spoofed audio detection method, namely Temporal Deepfake Location (TDL), which can effectively capt…

Cited by 0SourceScholar
2024

Binauralmusic: A Diverse Dataset for Improving Cross-Modal Binaural Audio Generation

ICASSP 2024accepted

Cross-modal binaural audio generation is an important task and has broad applications such as game sound development and auditory assistance for the visually impaired. However, existing datasets lack binaural samples with abundant visual venues. As a consequence, state-of-the-art cross-modal binaura…

Cited by 0SourceScholar
2024

DNIT: Enhancing Day-Night Image-to-Image Translation through Fine-Grained Feature Handling (Student Abstract)

AAAI 2024technical

Existing image-to-image translation methods perform less satisfactorily in the "day-night" domain due to insufficient scene feature study. To address this problem, we propose DNIT, which performs fine-grained handling of features by a nighttime image preprocessing (NIP) module and an edge fusion det…

Cited by 0SourcePDFScholar
2024

FSD: An Initial Chinese Dataset for Fake Song Detection

ICASSP 2024accepted

Singing voice synthesis and singing voice conversion have significantly advanced, revolutionizing musical experiences. However, the rise of "Deepfake Songs" generated by these technologies raises concerns about authenticity. Unlike Audio DeepFake Detection (ADD), the field of song deepfake detection…

Cited by 0SourceScholar