← Search

Yongzhi Li

7 accepted papers

2026

Adaptive Video Distillation: Mitigating Oversaturation and Temporal Collapse in Few-Step Generation

CVPR 2026

Video generation has recently emerged as a central task in the field of generative AI. However, the substantial computational cost inherent in video synthesis makes model distillation a critical technique for efficient deployment. Despite its significance, there is a scarcity of methods specifically

Cited by 0SourceScholar
2026

AutoCut: End-to-end advertisement video editing based on multimodal discretization and controllable generation

CVPR 2026

Short-form videos have become a primary medium for digital advertising, requiring scalable and efficient content creation. However, current workflows and AI tools remain disjoint and modality-specific, leading to high production costs and low overall efficiency. To address this issue, we propose Aut

Cited by 0SourcecodeScholar
2026

Narrative Weaver: Towards Controllable Long-Range Visual Consistency with Multi-Modal Conditioning

CVPR 2026

We present Narrative Weaver, a novel framework that addresses a fundamental challenge in generative AI: achieving controllable, long-range, and consistent visual content generation. While existing models excel at generating high-fidelity short-form visual content, they struggle to maintain narrative

Cited by 0SourcecodeScholar
2023

Learning Instance-Level Representation for Large-Scale Multi-Modal Pretraining in E-Commerce

CVPR 2023poster

This paper aims to establish a generic multi-modal foundation model that has the scalable capability to massive downstream applications in E-commerce. Recently, large-scale vision-language pretraining approaches have achieved remarkable advances in the general domain. However, due to the significant…

Cited by 13SourcePDFScholar
2022

Embracing Consistency: A One-Stage Approach for Spatio-Temporal Video Grounding

NeurIPS 2022accept

Spatio-Temporal video grounding (STVG) focuses on retrieving the spatio-temporal tube of a specific object depicted by a free-form textual expression. Existing approaches mainly treat this complicated task as a parallel frame-grounding problem and thus suffer from two types of inconsistency drawback…

2021

Learning 3-D Human Pose Estimation from Catadioptric Videos

IJCAI 2021poster

3-D human pose estimation is a crucial step for understanding human actions. However, reliably capturing precise 3-D position of human joints is non-trivial and tedious. Current models often suffer from the scarcity of high-quality 3-D annotated training data. In this work, we explore a novel way of…

Cited by 4SourcePDFScholar