← Search

Min Zhao

16 accepted papers

2026

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Video Generation

ICML 2026poster

To achieve real-time video generation, current approaches distill pretrained bidirectional video diffusion models into few-step autoregressive (AR) models. This process involves an *architectural gap*, as it converts full attention into causal attention. In this paper, we demonstrate that existing m…

Cited by 0SourceScholar
2026

SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse–Linear Attention

ICLR 2026poster

In Diffusion Transformer (DiT) models, particularly for video generation, attention latency is a major bottleneck due to the long sequence length and the quadratic complexity. Interestingly, we find that attention weights can be decoupled into two matrices: a small fraction of large weights with hig…

Cited by 44SourcecodeScholar
2026

UltraViCo: Breaking Extrapolation Limits in Video Diffusion Transformers

ICLR 2026poster

Despite advances, video diffusion transformers still struggle to generalize beyond their training length, a challenge we term video length extrapolation. We identify two failure modes: model-specific periodic content repetition and a universal quality degradation. Prior works attempt to solve repeti…

Cited by 0SourcecodeScholar
2025

FlexWorld: Progressively Expanding 3D Scenes for Flexible-View Exploration

NeurIPS 2025poster

Generating flexible-view 3D scenes, including 360° rotation and zooming, from single images is challenging due to a lack of 3D data. To this end, we introduce FlexWorld, a novel framework that progressively constructs a persistent 3D Gaussian splatting representation by synthesizing and integrating…

Cited by 0SourcecodeScholar
2025

RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion Transformers

ICML 2025poster

Recent advancements in video generation have enabled models to synthesize high-quality, minute-long videos. However, generating even longer videos with temporal coherence remains a major challenge and existing length extrapolation methods lead to temporal repetition or motion deceleration. In this w…

Cited by 0SourcePDFScholar
2024

Identifying and Solving Conditional Image Leakage in Image-to-Video Diffusion Model

NeurIPS 2024poster

Diffusion models have obtained substantial progress in image-to-video generation. However, in this paper, we find that these models tend to generate videos with less motion than expected. We attribute this to the issue called conditional image leakage, where the image-to-video diffusion models (I2V-…

2024

Modal Consensus and Contextual Separation for Weakly Supervised Temporal Action Localization

ICASSP 2024accepted

Weakly-supervised Temporal Action Localization (W-TAL) is a challenging task aiming to achieve both action class identification and localization of temporal boundaries using video-level label learning. Recent methods resort to basic cascading or integration of appearance and optical flow features, o…

Cited by 0SourceScholar
2024

PoseCrafter: One-Shot Personalized Video Synthesis Following Flexible Pose Control

ECCV 2024poster

"In this paper, we introduce PoseCrafter, a one-shot method for personalized video generation following the control of flexible poses. Built upon Stable Diffusion and ControlNet, we carefully design an inference process to produce high-quality videos without the corresponding ground-truth frames. Fi…

2023

Equivariant Energy-Guided SDE for Inverse Molecular Design

ICLR 2023poster

Inverse molecular design is critical in material science and drug discovery, where the generated molecules should satisfy certain desirable properties. In this paper, we propose equivariant energy-guided stochastic differential equations (EEGSDE), a flexible framework for controllable 3D molecule ge…

2022

AB-Mapper: Attention and BicNet based Multi-agent Path Planning for Dynamic Environment

IROS 2022poster

Multi-agent path finding in dynamic environments is of great academic and practical value for multi-robot systems in the real world. To improve the effectiveness and efficiency of the learning process during path planning in dynamic environments, we introduce an algorithm called Attention and BicNet…

Cited by 14SourceScholar
2022

CCRobot-V: A Silkworm-Like Cooperative Cable-Climbing Robotic System for Cable Inspection and Maintenance

ICRA 2022poster

This paper presents CCRobot-V, the fifth version of CCRobot, a cooperative serial multi-robot system for bridge cable inspection and maintenance that uses silkworm-like locomotion to climb the entire length of super-long stay cable at high speeds while carrying heavy inspection/maintenance equipment…

Cited by 13SourceScholar
2022

EGSDE: Unpaired Image-to-Image Translation via Energy-Guided Stochastic Differential Equations

NeurIPS 2022accept

Score-based diffusion models (SBDMs) have achieved the SOTA FID results in unpaired image-to-image translation (I2I). However, we notice that existing methods totally ignore the training data in the source domain, leading to sub-optimal solutions for unpaired I2I. To this end, we propose energy-gui…

2021

A General Framework for Lifelong Localization and Mapping in Changing Environment

IROS 2021poster

The environment of most real-world scenarios such as malls and supermarkets changes at all times. A pre-built map that does not account for these changes becomes out-of-date easily. Therefore, it is necessary to have an up-to-date model of the environment to facilitate long-term operation of a robot…

Cited by 46SourcecodeScholar
2021

CCRobot-IV-F: A Ducted-Fan-Driven Flying-Type Bridge-Stay-Cable Climbing Robot

IROS 2021poster

A Flying-type cable climbing robot, CCRobot-IV-F, is presented in this paper. It is a climbing precursor of the fourth version of CCRobot, designed to surpass the abilities of previous robots with high climbing speed and obstacle-crossing capability. CCRobot-IV-F weighs less than 10 kg and a no-load…

Cited by 10SourceScholar
2021

Proactive Interaction Framework for Intelligent Social Receptionist Robots

ICRA 2021poster

Proactive human-robot interaction (HRI) allows the receptionist robots to actively greet people and offer services based on vision, which has been found to improve acceptability and customer satisfaction. Existing approaches are either based on multi-stage decision processes or based on end-to-end d…

Cited by 13SourcecodeScholar
2021

Variational Autoencoders for Hyperspectral Unmixing with Endmember Variability

ICASSP 2021accepted

Spectral signatures are usually affected by variations in environmental conditions. The spectral variability is thus one of the most important and challenging problems to be addressed in hyperspectral unmixing. Generally, it is a non-trivial task to model the endmember variability, and existing spec…

Cited by 0SourceScholar