← Search

Yabo Chen

10 accepted papers

2026

SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video Generation

CVPR 2026

Controlling both camera motion and object dynamics is essential for coherent and expressive video generation, yet current methods typically handle only one motion type or rely on ambiguous 2D cues that entangle camera-induced parallax with true object movement. We present SymphoMotion, a unified mot

Cited by 1SourceScholar
2025

Diffusion-Driven Progressive Target Manipulation for Source-Free Domain Adaptation

NeurIPS 2025poster

Source-free domain adaptation (SFDA) is a challenging task that tackles domain shifts using only a pre-trained source model and unlabeled target data. Existing SFDA methods are restricted by the fundamental limitation of source-target domain discrepancy. Non-generation SFDA methods suffer from unrel…

Cited by 0SourceScholar
2025

IM-Zero: Instance-level Motion Controllable Video Generation in a Zero-shot Manner

CVPR 2025poster

Controllability of video generation has been recently concerned in addition to the quality of generated videos. The main challenge to controllable video generation is to synthesize videos based on user-specified instance spatial locations and movement trajectories. However, existing methods suffer f…

Cited by 0SourcePDFScholar
2024

Cascade-Zero123: One Image to Highly Consistent 3D with Self-Prompted Nearby Views

ECCV 2024poster

"Synthesizing multi-view 3D from one single image is a significant but challenging task. Zero-1-to-3 methods have achieved great success by lifting a 2D latent diffusion model to the 3D scope. The target-view image is generated with a single-view source image and the camera pose as condition informa…

2024

DomainFusion: Generalizing To Unseen Domains with Latent Diffusion Models

ECCV 2024poster

"Latent Diffusion Models (LDMs) are powerful and potential tools for facilitating generation-based methods for domain generalization. However, existing diffusion-based DG methods are restricted to offline augmentation using LDM and suffer from degraded performance and prohibitive computational costs…

Cited by 2SourcePDFScholar
2023

Progressively Compressed Auto-Encoder for Self-supervised Representation Learning

ICLR 2023poster

As a typical self-supervised learning strategy, Masked Image Modeling (MIM) is driven by recovering all masked patches from visible ones. However, patches from the same image are highly correlated and it is redundant to reconstruct all the masked patches. We find that this redundancy is neglected by…

2023

Promoting Semantic Connectivity: Dual Nearest Neighbors Contrastive Learning for Unsupervised Domain Generalization

CVPR 2023poster

Domain Generalization (DG) has achieved great success in generalizing knowledge from source domains to unseen target domains. However, current DG methods rely heavily on labeled source data, which are usually costly and unavailable. Since unlabeled data are far more accessible, we study a more pract…

Cited by 17SourcePDFScholar
2023

Towards Unsupervised Domain Generalization for Face Anti-Spoofing

ICCV 2023poster

Generalizable face anti-spoofing (FAS) based on domain generalization (DG) has gained growing attention due to its robustness in real-world applications. However, these DG methods rely heavily on labeled source data, which are usually costly and hard to access. Comparably, unlabeled face data are fa…

Cited by 43PDFScholar
2022

SdAE: Self-Distillated Masked Autoencoder

ECCV 2022poster

"With the development of generative-based self-supervised learning (SSL) approaches like BeiT and MAE, how to learn good representations by masking random patches of the input image and reconstructing the missing information has grown in concern. However, BeiT and PeCo need a “pre-pretraining” stage…

2022

Source-Free Domain Adaptation with Contrastive Domain Alignment and Self-Supervised Exploration for Face Anti-Spoofing

ECCV 2022poster

"Despite promising success in intra-dataset tests, existing face anti-spoofing (FAS) methods suffer from poor generalization ability under domain shift. This problem can be solved by aligning source and target data. However, due to privacy and security concerns of human faces, source data are usuall…