← Search

Long Ye

9 accepted papers

2026

Detect All-Type Deepfake Audio: Wavelet Prompt Tuning for Enhanced Auditory Perception

AAAI 2026technical

The rapid advancement of audio generation technologies has escalated the risks of malicious deepfake audio across speech, sound, singing voice, and music, threatening multimedia security and trust. While existing countermeasures (CMs) perform well in single-type audio deepfake detection (ADD), their

Cited by 0SourcePDFScholar
2026

MAG: Multi-Modal Aligned Autoregressive Co-Speech Gesture Generation without Vector Quantization

ICASSP 2026oral

This work focuses on full-body co-speech gesture generation. Existing methods typically employ an autoregressive model accompanied by vector-quantized tokens for gesture generation, which results in information loss and compromises the realism of the generated gestures. To address this, inspired by…

Cited by 0SourcePDFScholar
2025

Depth-Guided Bundle Sampling for Efficient Generalizable Neural Radiance Field Reconstruction

CVPR 2025poster

Recent advancements in generalizable novel view synthesis have achieved impressive quality through interpolation between nearby views. However, rendering high-resolution images remains computationally intensive due to the need for dense sampling of all rays. Recognizing that natural scenes are typic…

2025

GoLF-NRT: Integrating Global Context and Local Geometry for Few-Shot View Synthesis

CVPR 2025poster

Neural Radiance Fields (NeRF) have transformed novel view synthesis by modeling scene-specific volumetric representations directly from images. While generalizable NeRF models can generate novel views across unknown scenes by learning latent ray representations, their performance heavily depends on…

2024

An Efficient Temporary Deepfake Location Approach Based Embeddings for Partially Spoofed Audio Detection

ICASSP 2024accepted

Partially spoofed audio detection is a challenging task, lying in the need to accurately locate the authenticity of audio at the frame level. To address this issue, we propose a fine-grained partially spoofed audio detection method, namely Temporal Deepfake Location (TDL), which can effectively capt…

Cited by 0SourceScholar
2024

Binauralmusic: A Diverse Dataset for Improving Cross-Modal Binaural Audio Generation

ICASSP 2024accepted

Cross-modal binaural audio generation is an important task and has broad applications such as game sound development and auditory assistance for the visually impaired. However, existing datasets lack binaural samples with abundant visual venues. As a consequence, state-of-the-art cross-modal binaura…

Cited by 0SourceScholar
2024

DNIT: Enhancing Day-Night Image-to-Image Translation through Fine-Grained Feature Handling (Student Abstract)

AAAI 2024technical

Existing image-to-image translation methods perform less satisfactorily in the "day-night" domain due to insufficient scene feature study. To address this problem, we propose DNIT, which performs fine-grained handling of features by a nighttime image preprocessing (NIP) module and an edge fusion det…

Cited by 0SourcePDFScholar
2024

FSD: An Initial Chinese Dataset for Fake Song Detection

ICASSP 2024accepted

Singing voice synthesis and singing voice conversion have significantly advanced, revolutionizing musical experiences. However, the rise of "Deepfake Songs" generated by these technologies raises concerns about authenticity. Unlike Audio DeepFake Detection (ADD), the field of song deepfake detection…

Cited by 0SourceScholar