← Search

Zhenheng Yang

24 accepted papers

2026

Accelerating Autoregressive Video Diffusion via History-Guided Cache and Residual Correction

CVPR 2026

Caching-based acceleration methods have recently driven significant progress in efficient video generation with diffusion models. However, we identify a critical limitation when directly applying these acceleration techniques to auto-regressive video diffusion models, which generate long videos by s

Cited by 0SourceScholar
2026

FOCUS: Efficient Keyframe Selection for Long Video Understanding

ICLR 2026poster

Multimodal large language models (MLLMs) represent images and video frames as visual tokens. Scaling from single images to hour-long videos, however, inflates the token budget far beyond practical limits. Popular pipelines therefore either uniformly subsample or apply keyframe selection with retriev…

Cited by 0SourcecodeScholar
2026

Mixture of Contexts for Long Video Generation

ICLR 2026poster

Long video generation is fundamentally a long context memory problem: models must retain and retrieve salient events across a long range without collapsing or drifting. However, scaling diffusion transformers to generate long-context videos is fundamentally limited by the quadratic cost of self-atte…

Cited by 0SourceScholar
2026

MotionSight: Boosting Fine-Grained Motion Understanding in Multimodal LLMs

ICLR 2026poster

Despite advancements in Multimodal Large Language Models (MLLMs), their proficiency in fine-grained video motion understanding remains critically limited. They often lack inter-frame differencing and tend to average or ignore subtle visual cues. Furthermore, while visual prompting has shown potentia…

Cited by 0SourceScholar
2026

The Image as Its Own Reward: Reinforcement Learning with Adversarial Reward for Image Generation

CVPR 2026

A reliable reward function is essential for reinforcement learning (RL) in image generation. Most current RL approaches depend on pre-trained preference models that output scalar rewards to approximate human preferences. However, these rewards often fail to capture human perception and are vulnerabl

Cited by 0SourcecodeScholar
2026

UniAPO: Unified Multimodal Automated Prompt Optimization

AAAI 2026technical

Prompting is fundamental to unlocking the full potential of large language models. To automate and enhance this process, automatic prompt optimization (APO) has been developed, demonstrating effectiveness primarily in text-only input scenarios. However, extending existing APO methods to multimodal t

Cited by 0SourcePDFScholar
2026

VirtueBench: Evaluating Trustworthiness under Uncertainty in Long Video Understanding

CVPR 2026

Recent Vision-Language Models (VLMs) have made remarkable progress in multimodal understanding tasks, yet their evaluation on long video understanding remains unreliable. Due to limited frame inputs, key frames necessary for answering the question may be missing from the model's input. However, mode

Cited by 0SourceScholar
2025

DiCo: Revitalizing ConvNets for Scalable and Efficient Diffusion Modeling

NeurIPS 2025spotlight

Diffusion Transformer (DiT), a promising diffusion model for visual generation, demonstrates impressive performance but incurs significant computational overhead. Intriguingly, analysis of pre-trained DiT models reveals that global self-attention is often redundant, predominantly capturing local pat…

Cited by 0SourcecodeScholar
2025

InfiMM-WebMath-40B: Advancing Multimodal Pre-Training for Enhanced Mathematical Reasoning

EMNLP 2025

Pre-training on large, high-quality datasets is essential for improving the reasoning abilities of Large Language Models (LLMs), particularly in specialized fields like mathematics. However, the field of Multimodal LLMs (MLLMs) lacks a comprehensive, open-source dataset for mathematical reasoning. T

Cited by 0SourcePDFScholar
2025

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption

CVPR 2025poster

Text-to-video generation has evolved rapidly in recent years, delivering remarkable results. Training typically relies on video-caption paired data, which plays a crucial role in enhancing generation performance. However, current video captions often suffer from insufficient details, hallucinations…

2025

Long Context Tuning for Video Generation

ICCV 2025poster

Recent advances in video generation can produce realistic, minute-long single-shot videos with scalable diffusion transformers. However, real-world narrative videos require multi-shot scenes with visual and dynamic consistency across shots. In this work, we introduce Long Context Tuning (LCT), a tra…

Cited by 0SourcePDFScholar
2025

OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

ICLR 2025poster

Text-to-video (T2V) generation has recently garnered significant attention thanks to the large multi-modality model Sora. However, T2V generation still faces two important challenges: 1) Lacking a precise open sourced high-quality dataset. The previously popular video datasets, e.g.WebVid-10M and Pa…

Cited by 62SourcePDFScholar
2025

Parallelized Autoregressive Visual Generation

CVPR 2025highlight

Autoregressive models have emerged as a powerful approach for visual generation but suffer from slow inference speed due to their sequential token-by-token prediction process. In this paper, we propose a simple yet effective approach for parallelized autoregressive visual generation that improves ge…

2025

STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution

ICCV 2025poster

Image diffusion models have been adapted for real-world video super-resolution to tackle over-smoothing issues in GAN-based methods. However, these models struggle to maintain temporal consistency, as they are trained on static images, limiting their ability to capture temporal dynamics effectively.…

Cited by 0SourcePDFScholar
2025

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

ICLR 2025poster

We present a unified transformer, i.e., Show-o, that unifies multimodal understanding and generation. Unlike fully autoregressive models, Show-o unifies autoregressive and (discrete) diffusion modeling to adaptively handle inputs and outputs of various and mixed modalities. The unified model flexibl…

Cited by 164SourcePDFScholar
2021

Weakly Supervised Instance Segmentation for Videos With Temporal Mask Consistency

CVPR 2021poster

Weakly supervised instance segmentation reduces the cost of annotations required to train models. However, existing approaches which rely only on image-level class labels predominantly suffer from errors due to (a) partial segmentation of objects and (b) missing object predictions. We show that thes…

Cited by 31PDFScholar
2020

SPAN: Spatial Pyramid Attention Network for Image Manipulation Localization

ECCV 2020poster

Tehchniques for manipulating images are advancing rapidly; while these are helpful for many useful tasks, they also pose a threat to society with their ability to create believable misinformation. We present a novel, Spatial Pyramid Attention Network (SPAN) for detection and localization of multiple…

2019

Activity Driven Weakly Supervised Object Detection

CVPR 2019poster

Weakly supervised object detection aims at reducing the amount of supervision required to train detection models. Such models are traditionally learned from images/videos labelled only with the object class and not the object bounding box. In our work, we try to leverage not only the object class la…

Cited by 40PDFScholar
2019

UnOS: Unified Unsupervised Optical-Flow and Stereo-Depth Estimation by Watching Videos

CVPR 2019poster

In this paper, we propose UnOS, an unified system for unsupervised optical flow and stereo depth estimation using convolutional neural network (CNN) by taking advantages of their inherent geometrical consistency based on the rigid-scene assumption. UnOS significantly outperforms other state-of-the-a…

Cited by 194PDFScholar
2018

LEGO: Learning Edge With Geometry All at Once by Watching Videos

CVPR 2018poster

Learning to estimate 3D geometry in a single image by watching unlabeled videos via deep convolutional network is attracting significant attention. In this paper, we introduce a “3D as-smooth-as-possible (3D-ASAP)” prior inside the pipeline, which enables joint estimation of edges and 3D scene, yiel…

Cited by 215SourcePDFScholar
2018

Occlusion Aware Unsupervised Learning of Optical Flow

CVPR 2018poster

It has been recently shown that a convolutional neural network can learn optical flow estimation with unsuper- vised learning. However, the performance of the unsuper- vised methods still has a relatively large gap compared to its supervised counterpart. Occlusion and large motion are some of the ma…

Cited by 376SourcePDFScholar
2017

TURN TAP: Temporal Unit Regression Network for Temporal Action Proposals

ICCV 2017poster

We address the problem of Temporal Action Proposal (TAP) generation. This is an important problem, as fast extraction of semantically important (e.g. human actions) segments from untrimmed videos is an important step for large-scale video analysis. To tackle this problem, we propose a novel Temporal…

Cited by 486PDFcodeScholar