← Search

Huan Yang

22 accepted papers

2026

Mod-Adapter: Tuning-Free and Versatile Multi-concept Personalization via Modulation Adapter

ICLR 2026poster

Personalized text-to-image generation aims to synthesize images of user-provided concepts in diverse contexts. Despite recent progress in multi-concept personalization, most are limited to object concepts and struggle to customize abstract concepts (e.g., pose, lighting). Some methods have begun ex…

Cited by 0SourcecodeScholar
2026

SpatialDiff: 3D-Aware Object Movement via Implicit Spatial Modeling

CVPR 2026

Recent advances in image editing allow impressive manipulation of objects, existing methods still struggle to handle spatial movement in complex scenes, such as objects span different depth layers or are partially occluded. Most image editing methods focus solely on prior information from 2D dataset

Cited by 0SourceScholar
2026

TexEditor: Structure-Preserving Text-Driven texture Editing

ICML 2026poster

Text-guided texture editing aims to modify object appearance while preserving the underlying geometric structure. However, our empirical analysis reveals that even SOTA editing models frequently struggle to maintain structural consistency during texture editing, despite the intended changes being pu…

Cited by 0SourceScholar
2025

Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference Optimization

NeurIPS 2025poster

Preference optimization for diffusion models aims to align them with human preferences for images. Previous methods typically use Vision-Language Models (VLMs) as pixel-level reward models to approximate human preferences. However, when used for step-level preference optimization, these models face…

Cited by 0SourcecodeScholar
2025

Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation

CVPR 2025poster

We introduce Presto, a novel video diffusion model designed to generate 15-second videos with long-range coherence and rich content. Extending video generation to maintain scenario diversity over long durations presents significant challenges. To address this, we propose a Segmented Cross-Attention…

2025

Tele-GS: 3D Gaussian Scene Representation for Low-Bandwidth Teleoperation

IROS 2025

Video streaming based teleoperation often faces a trade-off between bandwidth consumption and the need for high-fidelity telepresence. Higher image resolution or a wider field of view (FOV) substantially increases bandwidth requirements. In this paper, we propose a novel telepresence model for teleo

Cited by 0SourceScholar
2025

VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation

CVPR 2025poster

Current large multimodal models (LMMs) face significant challenges in processing and comprehending long-duration or high-resolution videos, which is mainly due to the lack of high-quality datasets. To address this issue from a data-centric perspective, we propose VISTA, a simple yet effective video…

Cited by 4SourcePDFScholar
2025

Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers

ICCV 2025poster

State-of-the-art transformer-based large multimodal models (LMMs) struggle to handle hour-long video inputs due to the quadratic complexity of the causal self-attention operations, leading to high computational costs during training and inference. Existing token compression-based methods reduce the…

Cited by 0SourcePDFScholar
2024

Solving Diffusion ODEs with Optimal Boundary Conditions for Better Image Super-Resolution

ICLR 2024poster

Diffusion models, as a kind of powerful generative model, have given impressive results on image super-resolution (SR) tasks. However, due to the randomness introduced in the reverse process of diffusion models, the performances of diffusion-based SR models are fluctuating at every time of sampling,…

Cited by 9SourcePDFScholar
2024

Zero-Reference Low-Light Enhancement via Physical Quadruple Priors

CVPR 2024poster

Understanding illumination and reducing the need for supervision pose a significant challenge in low-light enhancement. Current approaches are highly sensitive to data usage during training and illumination-specific hyper-parameters limiting their ability to handle unseen scenarios. In this paper we…

2023

Learning Data-Driven Vector-Quantized Degradation Model for Animation Video Super-Resolution

ICCV 2023poster

Existing real-world video super-resolution (VSR) methods focus on designing a general degradation pipeline for open-domain videos while ignoring data intrinsic characteristics which strongly limit their performance when applying to some specific domains (e.g., animation videos). In this paper, we th…

Cited by 5PDFcodeScholar
2023

MM-Diffusion: Learning Multi-Modal Diffusion Models for Joint Audio and Video Generation

CVPR 2023poster

We propose the first joint audio-video generation framework that brings engaging watching and listening experiences simultaneously, towards high-quality realistic videos. To generate joint audio-video pairs, we propose a novel Multi-Modal Diffusion model (i.e., MM-Diffusion), with two-coupled denois…

2023

NUWA-XL: Diffusion over Diffusion for eXtremely Long Video Generation

ACL 2023long

In this paper, we propose NUWA-XL, a novel Diffusion over Diffusion architecture for eXtremely Long video generation. Most current work generates long videos segment by segment sequentially, which normally leads to the gap between training on short videos and inferring long videos, and the sequentia…

Cited by 118SourcePDFScholar
2022

Advancing High-Resolution Video-Language Representation With Large-Scale Video Transcriptions

CVPR 2022poster

We study joint video and language (VL) pre-training to enable cross-modality learning and benefit plentiful downstream VL tasks. Existing works either extract low-quality video features or learn limited text embedding, while neglecting that high-resolution videos and diversified semantics can signif…

Cited by 225PDFcodeScholar
2022

Learning Spatiotemporal Frequency-Transformer for Compressed Video Super-Resolution

ECCV 2022poster

"Compressed video super-resolution (VSR) aims to restore high-resolution frames from compressed low-resolution counterparts. Most recent VSR approaches often enhance an input frame by “borrowing’’ relevant textures from neighboring video frames. Although some progress has been made, there are grand…

2022

Long-Form Video-Language Pre-Training with Multimodal Temporal Contrastive Learning

NeurIPS 2022accept

Large-scale video-language pre-training has shown significant improvement in video-language understanding tasks. Previous studies of video-language pretraining mainly focus on short-form videos (i.e., within 30 seconds) and sentences, leaving long-form video-language pre-training rarely explored. Di…

2021

Improving Visual Quality of Image Synthesis by A Token-based Generator with Transformers

NeurIPS 2021poster

We present a new perspective of achieving image synthesis by viewing this task as a visual token generation problem. Different from existing paradigms that directly synthesize a full image from a single input (e.g., a latent code), the new formulation enables a flexible local manipulation for differ…

Cited by 33SourcePDFScholar
2021

Learning Conditional Knowledge Distillation for Degraded-Reference Image Quality Assessment

ICCV 2021poster

An important scenario for image quality assessment (IQA) is to evaluate image restoration (IR) algorithms. The state-of-the-art approaches adopt a full-reference paradigm that compares restored images with their corresponding pristine-quality images. However, pristine-quality images are usually unav…

Cited by 62PDFcodeScholar
2020

Learning Texture Transformer Network for Image Super-Resolution

CVPR 2020poster

We study on image super-resolution (SR), which aims to recover realistic textures from a low-resolution (LR) image. Recent progress has been made by taking high-resolution images as references (Ref), so that relevant textures can be transferred to LR images. However, existing SR approaches neglect t…

Cited by 1107PDFcodeScholar
2015

Unsupervised Extraction of Video Highlights Via Robust Recurrent Auto-Encoders

ICCV 2015poster

With the growing popularity of short-form video sharing platforms such as Instagram and Vine, there has been an increasing need for techniques that automatically extract highlights from video. Whereas prior works have approached this problem with heuristic rules or supervised learning, we present an…

Cited by 221PDFScholar