← Search

Zhengyao Lv

10 accepted papers

2026

DUO-VSR: Dual-Stream Distillation for One-Step Video Super-Resolution

CVPR 2026

Diffusion-based video super-resolution (VSR) achieves remarkable fidelity but suffers from prohibitive sampling cost. While distribution matching distillation (DMD) accelerates diffusion models to one-step generation, directly applying it to VSR leads to training instability and degraded, insufficie

Cited by 0SourcecodeScholar
2026

FlowSteer: Guiding Few-Step Image Synthesis with Authentic Trajectories

CVPR 2026

With the success of flow matching in visual generation, sampling efficiency remains a critical bottleneck for its practical application. Among flow models' accelerating methods, ReFlow has been somehow overlooked although it has theoretical consistency with flow matching. This is primarily due to it

Cited by 0SourceScholar
2026

NOVA: Sparse Control, Dense Synthesis for Pair-Free Video Editing

CVPR 2026

Recent video editing models have achieved impressive results, but most still require large-scale paired datasets. Collecting such naturally aligned pairs at scale remains highly challenging and constitutes a critical bottleneck, especially for local video editing data. Existing workarounds transfer

Cited by 0SourcecodeScholar
2025

Dual-Expert Consistency Model for Efficient and High-Quality Video Generation

ICCV 2025poster

Diffusion Models have achieved remarkable results in video synthesis but require iterative denoising steps, leading to substantial computational overhead. Consistency Models have made significant progress in accelerating diffusion models. However, directly applying them to video diffusion models oft…

2025

FasterCache: Training-Free Video Diffusion Model Acceleration with High Quality

ICLR 2025poster

In this paper, we present \textbf{\textit{FasterCache}}, a novel training-free strategy designed to accelerate the inference of video diffusion models with high-quality generation. By analyzing existing cache-based methods, we observe that \textit{directly reusing adjacent-step features degrades vid…

Cited by 6SourcePDFScholar
2025

Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers

ICCV 2025poster

Multimodal Diffusion Transformers (MM-DiTs) have achieved remarkable progress in text-driven visual generation. However, even state-of-the-art MM-DiT models like FLUX struggle with achieving precise alignment between text prompts and generated content. We identify two key issues in the attention mec…

2024

ConceptExpress: Harnessing Diffusion Models for Single-image Unsupervised Concept Extraction

ECCV 2024oral

"While personalized text-to-image generation has enabled the learning of a single concept from multiple images, a more practical yet challenging scenario involves learning multiple concepts within a single image. However, existing works tackling this scenario heavily rely on extensive human annotati…

2024

PLACE: Adaptive Layout-Semantic Fusion for Semantic Image Synthesis

CVPR 2024highlight

Recent advancements in large-scale pre-trained text-to-image models have led to remarkable progress in semantic image synthesis. Nevertheless synthesizing high-quality images with consistent semantics and layout remains a challenge. In this paper we propose the adaPtive LAyout-semantiC fusion modulE…

2022

Semantic-Shape Adaptive Feature Modulation for Semantic Image Synthesis

CVPR 2022poster

Recent years have witnessed substantial progress in semantic image synthesis, it is still challenging in synthesizing photo-realistic images with rich details. Most previous methods focus on exploiting the given semantic map, which just captures an object-level layout for an image. Obviously, a fine…

Cited by 35PDFcodeScholar
2021

Learning Semantic Person Image Generation by Region-Adaptive Normalization

CVPR 2021poster

Human pose transfer has received great attention due to its wide applications, yet is still a challenging task that is not well solved. Recent works have achieved great success to transfer the person image from the source to the target pose. However, most of them cannot well capture the semantic app…

Cited by 81PDFcodeScholar