← Search

Ye Zhu

30 accepted papers

2026

DAPointMamba: Domain Adaptive Point Mamba for Point Cloud Completion

AAAI 2026technical

Domain adaptive point cloud completion (DA PCC) aims to narrow the geometric and semantic discrepancies between the labeled source and unlabeled target domains. Existing methods either suffer from limited receptive fields or quadratic complexity due to using CNNs or vision Transformers. In this pape

Cited by 0SourcePDFScholar
2026

Forensic Prompting with Dual-Action Policy Optimization for Vision-Language Forgery Detection and Localization

ICML 2026poster

Image forgery is rapidly evolving, rendering forensic traces increasingly subtle and readily attenuated by post-processing. Although vision--language prompting can inject priors, open-ended LLM-generated prompts are difficult to constrain, and naive language description can introduce semantic pertur…

Cited by 0SourceScholar
2026

GASS: Geometry-Aware Spherical Sampling for Disentangled Diversity Enhancement in Text-to-Image Generation

ICML 2026poster

Despite high semantic alignment, modern text-to-image (T2I) generative models still struggle to synthesize diverse images from a given prompt. This lack of diversity not only restricts user choice, but also risks amplifying societal biases. In this work, we enhance the T2I diversity through a geomet…

Cited by 0SourceScholar
2026

Restoring Initial Noise Sensitivity in Text-to-Image Distillation through Geometric Alignment

ICML 2026poster

Generative distillation significantly accelerates text-to-image (T2I) generation by compressing multi-step trajectories into few-step student models while preserving perceptual quality. However, existing distillation methods prioritize efficiency and output fidelity, often overlooking the preservati…

Cited by 0SourceScholar
2026

SCoNE: Spherical Consistent Neighborhoods Ensemble for Effective and Efficient Multi-View Anomaly Detection

AAAI 2026technical

The core problem in multi-view anomaly detection is to represent local neighborhoods of normal instances consistently across all views. Recent approaches consider a representation of local neighborhood in each view independently, and then capture the consistent neighbors across all views via a learn

Cited by 0SourcePDFScholar
2025

BNMusic: Blending Environmental Noises into Personalized Music

NeurIPS 2025poster

While being disturbed by environmental noises, the acoustic masking technique is a conventional way to reduce the annoyance in audio engineering that seeks to cover up the noises with other dominant yet less intrusive sounds. However, misalignment between the dominant sound and the noise—such as mis…

Cited by 0SourcecodeScholar
2025

CanonSwap: High-Fidelity and Consistent Video Face Swapping via Canonical Space Modulation

ICCV 2025poster

Video face swapping aims to address two primary challenges: effectively transferring the source identity to the target video and accurately preserving the dynamic attributes of the target face, such as head poses, facial expressions, lip-sync, etc..Existing methods mainly focus on achieving high-qua…

Cited by 0SourcePDFScholar
2025

DAPoinTr: Domain Adaptive Point Transformer for Point Cloud Completion

AAAI 2025technical

Point Transformers (PoinTr) have shown great potential in point cloud completion recently. Nevertheless, effective domain adaptation that improves transferability toward target domains remains unexplored. In this paper, we delve into this topic and empirically discover that direct feature alignment…

2025

D^3: Scaling Up Deepfake Detection by Learning from Discrepancy

CVPR 2025poster

The boom of Generative AI brings opportunities entangled with risks and concerns. Existing literature emphasizes the generalization capability of deepfake detection on unseen generators, significantly promoting the detector's ability to identify more universal artifacts. This work seeks a step towar…

2025

Dynamic Diffusion Schrödinger Bridge in Astrophysical Observational Inversions

NeurIPS 2025poster

We study Diffusion Schrödinger Bridge (DSB) models in the context of dynamical astrophysical systems, specifically tackling observational inverse prediction tasks within Giant Molecular Clouds (GMCs) for star formation. We introduce the Astro-DSB model, a variant of DSB with the pairwise domain assu…

Cited by 0SourcecodeScholar
2025

GUAVA: Generalizable Upper Body 3D Gaussian Avatar

ICCV 2025poster

Reconstructing a high-quality, animatable 3D human avatar with expressive facial and hand motions from a single image has gained significant attention due to its broad application potential. 3D human avatar reconstruction typically requires multi-view or monocular videos and training on individual I…

Cited by 0SourcePDFScholar
2025

HRAvatar: High-Quality and Relightable Gaussian Head Avatar

CVPR 2025poster

Reconstructing animatable and high-quality 3D head avatars from monocular videos, especially with realistic relighting, is a valuable task. However, the limited information from single-view input, combined with the complex head poses and facial movements, makes this challenging. Previous methods ach…

Cited by 0SourcePDFScholar
2025

TEASER: Token Enhanced Spatial Modeling for Expressions Reconstruction

ICLR 2025poster

3D facial reconstruction from a single in-the-wild image is a crucial task in human-centered computer vision tasks. While existing methods can recover accurate facial shapes, there remains significant space for improvement in fine-grained expression capture. Current approaches struggle with irregul…

Cited by 1SourcePDFScholar
2025

The Silent Assistant: NoiseQuery as Implicit Guidance for Goal-Driven Image Generation

ICCV 2025poster

In this work, we introduce NoiseQuery as a novel method for enhanced noise initialization in versatile goal-driven text-to-image (T2I) generation. Specifically, we propose to leverage an aligned Gaussian noise as implicit guidance to complement explicit user-defined inputs, such as text prompts, for…

2024

A Local-Ascending-Global Learning Strategy for Brain-Computer Interface

AAAI 2024technical

Neuroscience research indicates that the interaction among different functional regions of the brain plays a crucial role in driving various cognitive tasks. Existing studies have primarily focused on constructing either local or global functional connectivity maps within the brain, often lacking an…

Cited by 9SourcePDFScholar
2024

Detecting Change Intervalswith Isolation Distributional Kernel (Abstract Reprint)

IJCAI 2024poster

Detecting abrupt changes in data distribution is one of the most significant tasks in streaming data analysis. Although many unsupervised Change-Point Detection (CPD) methods have been proposed recently to identify those changes, they still suffer from missing subtle changes, poor scalability, or/an…

Cited by 0SourcePDFScholar
2024

Diffusion in Diffusion: Cyclic One-Way Diffusion for Text-Vision-Conditioned Generation

ICLR 2024poster

Originating from the diffusion phenomenon in physics that describes particle movement, the diffusion generative models inherit the characteristics of stochastic random walk in the data space along the denoising trajectory. However, the intrinsic mutual interference among image regions contradicts th…

2024

Supplementing Missing Visions Via Dialog for Scene Graph Generations

ICASSP 2024accepted

Most AI systems rely on the premise that the input visual data are sufficient to achieve competitive performance in various tasks. However, the classic task setup rarely considers the challenging, yet common practical situations where the complete visual data may be inaccessible due to various reaso…

Cited by 0SourceScholar
2023

Boundary Guided Learning-Free Semantic Control with Diffusion Models

NeurIPS 2023poster

Applying pre-trained generative denoising diffusion models (DDMs) for downstream tasks such as image semantic editing usually requires either fine-tuning DDMs or learning auxiliary editing networks in the existing literature. In this work, we present our BoundaryDiffusion method for efficient, effec…

2023

Discrete Contrastive Diffusion for Cross-Modal Music and Image Generation

ICLR 2023poster

Diffusion probabilistic models (DPMs) have become a popular approach to conditional generation, due to their promising results and support for cross-modal synthesis. A key desideratum in conditional synthesis is to achieve high correspondence between the conditioning input and generated output. Most…

2023

Towards a Persistence Diagram that is Robust to Noise and Varied Densities

ICML 2023poster

Recent works have identified that existing methods, which construct persistence diagrams in Topological Data Analysis (TDA), are not robust to noise and varied densities in a point cloud. We analyze the necessary properties of an approach that can address these two issues, and propose a new filter f…

Cited by 1SourcePDFScholar
2022

AMOS: A Large-Scale Abdominal Multi-Organ Benchmark for Versatile Medical Image Segmentation

NeurIPS 2022accept

Despite the considerable progress in automatic abdominal multi-organ segmentation from CT/MRI scans in recent years, a comprehensive evaluation of the models' capabilities is hampered by the lack of a large-scale benchmark from diverse clinical scenarios. Constraint by the high cost of collecting an…

2022

Improving the Effectiveness and Efficiency of Stochastic Neighbour Embedding with Isolation Kernel (Extended Abstract)

IJCAI 2022poster

This paper presents a new insight into improving the performance of Stochastic Neighbour Embedding (t-SNE) by using Isolation kernel instead of Gaussian kernel. We show that Isolation kernel addresses two deficiencies of t-SNE that employs Gaussian kernel, and the use of Isolation kernel enables t-S…

2022

Quantized GAN for Complex Music Generation from Dance Videos

ECCV 2022poster

"We present Dance2Music-GAN (D2M-GAN), a novel adversarial multi-modal framework that generates complex musical samples conditioned on dance videos. Our proposed framework takes dance video frames and human body motions as input, and learns to generate music samples that plausibly accompany the corr…

2021

Learning Audio-Visual Correlations From Variational Cross-Modal Generation

ICASSP 2021accepted

People can easily imagine the potential sound while seeing an event. This natural synchronization between audio and visual signals reveals their intrinsic correlations. To this end, we propose to learn the audio-visual correlations from the perspective of cross-modal generation in a self-supervised…

Cited by 0SourceScholar