← Search

Zhaofan Qiu

35 accepted papers

2026

Distillation Models are Good Samplers for Diffusion Reinforcement Learning

ICML 2026poster

We present DMSampler, a framework that accelerates diffusion reinforcement learning by using fast distillation models as its training-time sampling engine. It overcomes the key bottleneck of sampling from the policy model—typically requiring around 50 denoising steps—by employing a co-evolving disti…

Cited by 0SourceScholar
2026

EvoID: Reinforced Evolution for Identity-Preserving Video Generation

CVPR 2026

We present EvoID, a novel framework that reformulates Identity-Preserving Video Generation as a self-evolving process through Reinforcement Learning. Moving beyond the static paradigm of imitation learning, EvoID enables a generative model to actively learn and optimize the complex trade-offs betwee

Cited by 0SourceScholar
2026

In-Context Generation with Regional Constraints for Instructional Video Editing

ICML 2026poster

The In-context generation paradigm has demonstrated strong power in instructional image editing for better synthesis quality. Nevertheless, shaping such in-context learning for instructional video editing is not trivial. Without specifying editing regions, the results can suffer from the issue of in…

Cited by 0SourceScholar
2026

ReactID: Synchronizing Realistic Actions and Identity in Personalized Video Generation

ICLR 2026poster

Personalized video generation faces a fundamental trade-off between identity consistency and action realism: overly rigid identity preservation often leads to unnatural motion, while emphasis on action dynamics can compromise subject fidelity. This tension stems from three interrelated challenges: i…

Cited by 0SourceScholar
2025

Aligning Global Semantics and Local Textures in Generative Video Enhancement

ICCV 2025poster

Recent advances in video generation have demonstrated the utility of powerful diffusion models. One important direction among them is to enhance the visual quality of the AI-synthesized videos for artistic creation. Nevertheless, solely relying on the knowledge embedded in the pre-trained video diff…

2025

MotionPro: A Precise Motion Controller for Image-to-Video Generation

CVPR 2025poster

Animating images with interactive motion control has garnered popularity for image-to-video (I2V) generation. Modern approaches typically rely on large Gaussian kernels to extend motion trajectories as condition without explicitly defining movement region, leading to coarse motion control and failin…

2025

Ouroboros-Diffusion: Exploring Consistent Content Generation in Tuning-free Long Video Diffusion

AAAI 2025technical

The first-in-first-out (FIFO) video diffusion, built on a pre-trained text-to-video model, has recently emerged as an effective approach for tuning-free long video generation. This technique maintains a queue of video frames with progressively increasing noise, continuously producing clean frames at…

Cited by 0SourcePDFScholar
2024

Learning Spatial Adaptation and Temporal Coherence in Diffusion Models for Video Super-Resolution

CVPR 2024poster

Diffusion models are just at a tipping point for image super-resolution task. Nevertheless it is not trivial to capitalize on diffusion models for video super-resolution which necessitates not only the preservation of visual appearance from low-resolution to high-resolution videos but also the tempo…

Cited by 7SourcePDFScholar
2024

TRIP: Temporal Residual Learning with Image Noise Prior for Image-to-Video Diffusion Models

CVPR 2024poster

Recent advances in text-to-video generation have demonstrated the utility of powerful diffusion models. Nevertheless the problem is not trivial when shaping diffusion models to animate static image (i.e. image-to-video generation). The difficulty originates from the aspect that the diffusion process…

2023

AnchorFormer: Point Cloud Completion From Discriminative Nodes

CVPR 2023poster

Point cloud completion aims to recover the completed 3D shape of an object from its partial observation. A common strategy is to encode the observed points to a global feature vector and then predict the complete points through a generative process on this vector. Nevertheless, the results may suffe…

2023

Learning Neural Implicit Surfaces with Object-Aware Radiance Fields

ICCV 2023poster

Recent progress on multi-view 3D object reconstruction has featured neural implicit surfaces via learning high-fidelity radiance fields. However, most approaches hinge on the visual hull derived from cost-expensive silhouette masks to obtain object surfaces. In this paper, we propose a novel Object-…

Cited by 2PDFScholar
2023

Learning Orthogonal Prototypes for Generalized Few-Shot Semantic Segmentation

CVPR 2023poster

Generalized few-shot semantic segmentation (GFSS) distinguishes pixels of base and novel classes from the background simultaneously, conditioning on sufficient data of base classes and a few examples from novel class. A typical GFSS approach has two training phases: base class learning and novel cla…

2023

PointClustering: Unsupervised Point Cloud Pre-Training Using Transformation Invariance in Clustering

CVPR 2023highlight

Feature invariance under different data transformations, i.e., transformation invariance, can be regarded as a type of self-supervision for representation learning. In this paper, we present PointClustering, a new unsupervised representation learning scheme that leverages transformation invariance f…

2022

Dynamic Temporal Filtering In Video Models

ECCV 2022poster

"Video temporal dynamics is conventionally modeled with 3D spatial-temporal kernel or its factorized version comprised of 2D spatial kernel and 1D temporal kernel. The modeling power, nevertheless, is limited by the fixed window size and static weights of a kernel along the temporal dimension. The p…

2022

SPE-Net: Boosting Point Cloud Analysis via Rotation Robustness Enhancement

ECCV 2022poster

"In this paper, we propose a novel deep architecture tailored for 3D point cloud applications, named as SPE-Net. The embedded ""Selective Position Encoding (SPE)"" procedure relies on an attention mechanism that can effectively attend to the underlying rotation condition of the input. Such encoded r…

2022

Stand-Alone Inter-Frame Attention in Video Models

CVPR 2022poster

Motion, as the uniqueness of a video, has been critical to the development of video understanding models. Modern deep learning models leverage motion by either executing spatio-temporal 3D convolutions, factorizing 3D convolutions into spatial and temporal convolutions separately, or computing self-…

Cited by 62PDFcodeScholar
2021

Boosting Video Representation Learning With Multi-Faceted Integration

CVPR 2021poster

Video content is multifaceted, consisting of objects, scenes, interactions or actions. The existing datasets mostly label only one of the facets for model training, resulting in the video representation that biases to only one facet depending on the training dataset. There is no study yet on how to…

Cited by 13PDFScholar
2021

Condensing a Sequence to One Informative Frame for Video Recognition

ICCV 2021poster

Video is complex due to large variations in motion and rich content in fine-grained visual details. Abstracting useful information from such information-intensive media requires exhaustive computing resources. This paper studies a two-step alternative that first condenses the video sequence to an in…

Cited by 10PDFScholar
2021

Motion-Focused Contrastive Learning of Video Representations

ICCV 2021poster

Motion, as the most distinct phenomenon in a video to involve the changes over time, has been unique and critical to the development of video representation learning. In this paper, we ask the question: how important is the motion particularly for self-supervised video representation learning. To th…

Cited by 47PDFcodeScholar
2021

Representing Videos As Discriminative Sub-Graphs for Action Recognition

CVPR 2021poster

Human actions are typically of combinatorial structures or patterns, i.e., subjects, objects, plus spatio-temporal interactions in between. Discovering such structures is therefore a rewarding way to reason about the dynamics of interactions and recognize the actions. In this paper, we introduce a n…

Cited by 34PDFScholar
2021

SeCo: Exploring Sequence Supervision for Unsupervised Representation Learning

AAAI 2021technical

A steady momentum of innovations and breakthroughs has convincingly pushed the limits of unsupervised image representation learning. Compared to static 2D images, video has one more dimension (time). The inherent supervision existing in such sequential structure offers a fertile ground for building…

2020

Learning to Localize Actions from Moments

ECCV 2020poster

With the knowledge of action moments (i.e., trimmed video clips that each contains an action instance), humans could routinely localize an action temporally in an untrimmed video. Nevertheless, most practical methods still require all training videos to be labeled with temporal annotations (action c…

2020

Transferring and Regularizing Prediction for Semantic Segmentation

CVPR 2020poster

Semantic segmentation often requires a large set of images with pixel-level annotations. In the view of extremely expensive expert labeling, recent research has shown that the models trained on photo-realistic synthetic data (e.g., computer games) with computer-generated annotations can be adapted t…

Cited by 47PDFScholar
2019

Customizable Architecture Search for Semantic Segmentation

CVPR 2019poster

In this paper, we propose a Customizable Architecture Search (CAS) approach to automatically generate a network architecture for semantic image segmentation. The generated network consists of a sequence of stacked computation cells. A computation cell is represented as a directed acyclic graph, in w…

Cited by 179PDFScholar
2019

Gaussian Temporal Awareness Networks for Action Localization

CVPR 2019oral

Temporally localizing actions in a video is a fundamental challenge in video understanding. Most existing approaches have often drawn inspiration from image object detection and extended the advances, e.g., SSD and Faster R-CNN, to produce temporal locations of an action in a 1D sequence. Neverthele…

Cited by 438PDFScholar
2019

Learning Spatio-Temporal Representation With Local and Global Diffusion

CVPR 2019poster

Convolutional Neural Networks (CNN) have been regarded as a powerful class of models for visual recognition problems. Nevertheless, the convolutional filters in these networks are local operations while ignoring the large-range dependency. Such drawback becomes even worse particularly for video reco…

Cited by 236PDFScholar
2018

Fully Convolutional Adaptation Networks for Semantic Segmentation

CVPR 2018poster

The recent advances in deep neural networks have convincingly demonstrated high capability in learning vision models on large datasets. Nevertheless, collecting expert labeled datasets especially with pixel-level annotations is an extremely expensive process. An appealing alternative is to render sy…

Cited by 429SourcePDFScholar
2018

Recurrent Tubelet Proposal and Recognition Networks for Action Detection

ECCV 2018poster

Detecting actions in videos is a challenging task as video is an information intensive media with complex variations. Existing approaches predominantly generate action proposals for each individual frame or fixed-length clip independently, while overlooking temporal context across them. Such tempora…

Cited by 146SourcePDFScholar