← Search

Feng Cheng

12 accepted papers

2026

VINCIE: Unlocking In-context Image Editing from Video

ICLR 2026poster

In-context image editing aims to modify images based on a contextual sequence comprising text and previously generated images. Existing methods typically depend on task-specific pipelines and expert models (e.g., segmentation and inpainting) to curate training data. In this work, we explore whether…

Cited by 0SourcecodeScholar
2025

Synthetic Video Enhances Physical Fidelity in Video Synthesis

ICCV 2025poster

We investigate how to enhance the physical fidelity of video generation models by leveraging synthetic videos generated via standard computer graphics techniques. These rendered videos respect real-world physics -- such as maintaining 3D consistency -- thereby serving as a valuable resource that can…

2025

VideoAuteur: Towards Long Narrative Video Generation

ICCV 2025poster

Recent video generation models have shown promising results in producing high-quality video clips lasting several seconds. However, these models face challenges in generating long sequences that convey clear and informative events, limiting their ability to support coherent narrations. In this paper…

Cited by 0SourcePDFScholar
2025

VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

CVPR 2025poster

Long-form video understanding has been a challenging task due to the high redundancy in video data and the abundance of query-irrelevant information. To tackle this challenge, we propose VideoTree, a training-free framework which builds a query-adaptive and hierarchical video representation for LLM…

2024

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

CVPR 2024poster

We present Ego-Exo4D a diverse large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g. sports music dance bike repair). 740 participants from 13 cities worldwide perform…

2023

Unified Coarse-to-Fine Alignment for Video-Text Retrieval

ICCV 2023poster

The canonical approach to video-text retrieval leverages a coarse-grained or fine-grained alignment between visual and textual information. However, retrieving the correct video according to the text query is often challenging as it requires the ability to reason about both high-level (scene) and lo…

Cited by 58PDFcodeScholar
2023

VindLU: A Recipe for Effective Video-and-Language Pretraining

CVPR 2023poster

The last several years have witnessed remarkable progress in video-and-language (VidL) understanding. However, most modern VidL approaches use complex and specialized model architectures and sophisticated pretraining protocols, making the reproducibility, analysis and comparisons of these frameworks…

2022

Stochastic Backpropagation: A Memory Efficient Strategy for Training Video Models

CVPR 2022oral

We propose a memory efficient method, named Stochastic Backpropagation (SBP), for training deep neural networks on videos. It is based on the finding that gradients from incomplete execution for backpropagation can still effectively train the models with minimal accuracy loss, which attributes to th…

Cited by 22PDFcodeScholar
2017

Faster and Non-ergodic O(1/K) Stochastic Alternating Direction Method of Multipliers

NeurIPS 2017poster

We study stochastic convex optimization subjected to linear equality constraints. Traditional Stochastic Alternating Direction Method of Multipliers and its Nesterov's acceleration scheme can only achieve ergodic O(1/\sqrt{K}) convergence rates, where K is the number of iteration. By introducing Var…

Cited by 14SourcePDFScholar