← Search

Duo Peng

10 accepted papers

2026

Training-free Motion Factorization for Compositional Video Generation

CVPR 2026

Compositional video generation aims to synthesize multiple instances with diverse appearance and motion. However, current approaches mainly focus on binding semantics, neglecting to understand diverse motion categories specified in prompts. In this paper, we propose a motion factorization framework

Cited by 0SourcecodeScholar
2025

Training-free Dense-Aligned Diffusion Guidance for Modular Conditional Image Synthesis

CVPR 2025poster

Conditional image synthesis is a crucial task with broad applications, such as artistic creation and virtual reality. However, current generative methods are often task-oriented with a narrow scope, handling a restricted condition with constrained applicability. In this paper, we propose a novel app…

2025

Visual Prompting for One-shot Controllable Video Editing without Inversion

CVPR 2025poster

One-shot controllable video editing (OCVE) is an important yet challenging task, aiming to propagate user edits that are made---using any image editing tool---on the first frame of a video to all subsequent frames, while ensuring content consistency between edited frames and source frames. To achiev…

2024

Diff-Tracker: Text-to-Image Diffusion Models are Unsupervised Trackers

ECCV 2024poster

"We introduce Diff-Tracker, a novel approach for the challenging unsupervised visual tracking task leveraging the pre-trained text-to-image diffusion model. Our main idea is to leverage the rich knowledge encapsulated within the pre-trained diffusion model, such as the understanding of image semanti…

Cited by 12SourcePDFScholar
2024

Harnessing Text-to-Image Diffusion Models for Category-Agnostic Pose Estimation

ECCV 2024oral

"Category-Agnostic Pose Estimation (CAPE) aims to detect keypoints of an arbitrary unseen category in images, based on several provided examples of that category. This is a challenging task, as the limited data of unseen categories makes it difficult for models to generalize effectively. To address…

Cited by 10SourcePDFScholar
2024

UPAM: Unified Prompt Attack in Text-to-Image Generation Models Against Both Textual Filters and Visual Checkers

ICML 2024poster

Text-to-Image (T2I) models have raised security concerns due to their potential to generate inappropriate or harmful images. In this paper, we propose UPAM, a novel framework that investigates the robustness of T2I models from the attack perspective. Unlike most existing attack methods that focus on…

Cited by 4SourcePDFScholar
2023

Diffusion-based Image Translation with Label Guidance for Domain Adaptive Semantic Segmentation

ICCV 2023poster

Translating images from a source domain to a target domain for learning target models is one of the most common strategies in domain adaptive semantic segmentation (DASS). However, existing methods still struggle to preserve semantically-consistent local details between the original and translated i…

Cited by 32PDFScholar
2023

Joint Attribute and Model Generalization Learning for Privacy-Preserving Action Recognition

NeurIPS 2023poster

Privacy-Preserving Action Recognition (PPAR) aims to transform raw videos into anonymous ones to prevent privacy leakage while maintaining action clues, which is an increasingly important problem in intelligent vision applications. Despite recent efforts in this task, it is still challenging to deal…

Cited by 4SourcePDFScholar
2021

Sparse-to-Dense Feature Matching: Intra and Inter Domain Cross-Modal Learning in Domain Adaptation for 3D Semantic Segmentation

ICCV 2021poster

Domain adaptation is critical for success when confronting with the lack of annotations in a new domain. As the huge time consumption of labeling process on 3D point cloud, domain adaptation for 3D semantic segmentation is of great expectation. With the rise of multi-modal datasets, large amount of…

Cited by 66PDFcodeScholar