← Search

Wenhan Luo

51 accepted papers

2026

Beyond VLM-Based Rewards: Diffusion-Native Latent Reward Modeling

ICML 2026poster

Preference optimization for diffusion models relies on reward functions that are both discriminative and computationally efficient. Vision-Language Models (VLMs) have emerged as powerful reward providers. However, their computation and memory cost can be substantial, and optimizing a latent diffusio…

Cited by 0SourceScholar
2026

CogniEdit: Dense Gradient Flow Optimization for Fine-Grained Image Editing

CVPR 2026

Instruction-based image editing with diffusion models has achieved impressive results, yet existing methods struggle with fine-grained instructions specifying precise attributes such as colors, positions, and quantities. While recent approaches employ Group Relative Policy Optimization (GRPO) for al

Cited by 0SourcecodeScholar
2026

FlowSteer: Guiding Few-Step Image Synthesis with Authentic Trajectories

CVPR 2026

With the success of flow matching in visual generation, sampling efficiency remains a critical bottleneck for its practical application. Among flow models' accelerating methods, ReFlow has been somehow overlooked although it has theoretical consistency with flow matching. This is primarily due to it

Cited by 0SourceScholar
2026

MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding

CVPR 2026

Understanding high-resolution (HR) images remains a critical challenge for multimodal large language models (MLLMs). Recent approaches leverage vision-based retrieval-augmented generation (RAG) to retrieve query-relevant crops from HR images, improving understanding capacity of MLLMs. However, this

Cited by 0SourcecodeScholar
2026

Pixel-Perfect Puppetry: Precision-Guided Enhancement for Face Image and Video Editing

ICLR 2026poster

Preserving identity while precisely manipulating attributes is a central challenge in face editing for both images and videos. Existing methods often introduce visual artifacts or fail to maintain temporal consistency. We present **FlowGuide**, a unified framework that achieves fine-grained control…

Cited by 0SourceScholar
2026

STiTch: Semantic Transition and Transportation in Collaboration for Training-Free Zero-Shot Composed Image Retrieval

CVPR 2026

Training-free zero-shot composed image retrieval models are recently gaining increasing research interest due to their generalizability and flexibility in unseen multimodal retrieval. Recent LLM-based advances focus on generating the expected target caption by exploring the compositional ability beh

Cited by 0SourceScholar
2026

UniSH: Unifying Scene and Human Reconstruction in a Feed-Forward Pass

CVPR 2026

We present UniSH, a unified, feed-forward framework for joint metric-scale 3D scene and human reconstruction. A key challenge in this domain is the scarcity of large-scale, annotated real-world data, forcing a reliance on synthetic datasets. This reliance introduces a significant sim-to-real domain

Cited by 0SourcecodeScholar
2026

Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Models

CVPR 2026

Recently, the introduction of Chain-of-Thought (CoT) has largely improved generation ability of unified models. However, it is observed that the current thinking process during generation mainly focuses on the text consistency with the text prompt, ignoring the visual context consistency with the vi

Cited by 0SourceScholar
2025

Co$^{\mathbf{3}}$Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion

ICLR 2025spotlight

Generating gestures from human speech has gained tremendous progress in animating virtual avatars. While the existing methods enable synthesizing gestures cooperated by people self-talking, they overlook the practicality of concurrent gesture modeling with two-person interactive conversations. Moreo…

Cited by 0SourcePDFScholar
2025

Foundation Cures Personalization: Improving Personalized Models’ Prompt Consistency via Hidden Foundation Knowledge

NeurIPS 2025poster

Facial personalization faces challenges to maintain identity fidelity without disrupting the foundation model's prompt consistency. The mainstream personalization models employ identity embedding to integrate identity information within the attention mechanisms. However, our preliminary findings rev…

Cited by 0SourceScholar
2025

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

NeurIPS 2025poster

Audio-driven human animation methods, such as talking head and talking body generation, have made remarkable progress in generating synchronized facial movements and appealing visual quality videos. However, existing methods primarily focus on single human animation and struggle with multi-stream au…

Cited by 0SourcecodeScholar
2025

MOERL: When Mixture-of-Experts Meet Reinforcement Learning for Adverse Weather Image Restoration

ICCV 2025poster

Adverse weather conditions, such as rain, snow, and haze, introduce complex degradations that present substantial challenges for effective image restoration. Existing all-in-one models often rely on fixed network structures, limiting their ability to adapt to the varying characteristics of different…

Cited by 0SourcePDFScholar
2025

MaterialMVP: Illumination-Invariant Material Generation via Multi-view PBR Diffusion

ICCV 2025poster

Physically-based rendering (PBR) has become a cornerstone in modern computer graphics, enabling realistic material representation and lighting interactions in 3D scenes. In this paper, we present MaterialMVP, a novel end-to-end model for generating PBR textures from 3D meshes and image prompts, addr…

Cited by 0SourcePDFScholar
2025

OSV: One Step is Enough for High-Quality Image to Video Generation

CVPR 2025poster

Video diffusion models have shown great potential in generating high-quality videos, making them an increasingly popular focus. However, their inherent iterative nature leads to substantial computational and time costs. Although techniques such as consistency distillation and adversarial training ha…

Cited by 10SourcePDFScholar
2025

PSHuman: Photorealistic Single-image 3D Human Reconstruction using Cross-Scale Multiview Diffusion and Explicit Remeshing

CVPR 2025poster

Photorealistic 3D human modeling is essential for various applications and has seen tremendous progress. However, existing methods for monocular full-body reconstruction, typically relying on front and/or predicted back view, still struggle with satisfactory performance due to the ill-posed nature o…

2025

StyleMaster: Stylize Your Video with Artistic Generation and Translation

CVPR 2025poster

Style control has been popular in video generation models. Existing methods often generate videos far from the given style, cause content leakage, and struggle to transfer one video to the desired style. Our first observation is that the style extraction stage matters, whereas existing methods empha…

Cited by 3SourcePDFScholar
2025

Towards Multiple Character Image Animation Through Enhancing Implicit Decoupling

ICLR 2025poster

Controllable character image animation has a wide range of applications. Although existing studies have consistently improved performance, challenges persist in the field of character image animation, particularly concerning stability in complex backgrounds and tasks involving multiple characters. T…

Cited by 0SourcePDFScholar
2025

VideoVista-CulturalLingo: 360° Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension

ACL 2025long

Assessing the video comprehension capabilities of multimodal AI systems can effectively measure their understanding and reasoning abilities. Most video evaluation benchmarks are limited to a single language, typically English, and predominantly feature videos rooted in Western cultural contexts. In…

Cited by 0SourcePDFScholar
2024

A Multimodal In-Context Tuning Approach for E-Commerce Product Description Generation

COLING 2024main

In this paper, we propose a new setting for generating product descriptions from images, augmented by marketing keywords. It leverages the combined power of visual and textual information to create descriptions that are more tailored to the unique features of products. For this setting, previous met…

2024

AttnZero: Efficient Attention Discovery for Vision Transformers

ECCV 2024poster

"In this paper, we present AttnZero, the first framework for automatically discovering efficient attention modules tailored for Vision Transformers (ViTs). While traditional self-attention in ViTs suffers from quadratic computation complexity, linear attention offers a more efficient alternative wit…

2024

Auto-GAS: Automated Proxy Discovery for Training-free Generative Architecture Search

ECCV 2024poster

"In this paper, we introduce Auto-GAS, the first training-free Generative Architecture Search (GAS) framework enabled by an auto-discovered proxy. Generative models like Generative Adversarial Networks (GANs) are now widely used in many real-time applications. Previous GAS methods use differentiable…

2024

Aux-NAS: Exploiting Auxiliary Labels with Negligibly Extra Inference Cost

ICLR 2024poster

We aim at exploiting additional auxiliary labels from an independent (auxiliary) task to boost the primary task performance which we focus on, while preserving a single task inference cost of the primary task. While most existing auxiliary learning methods are optimization-based relying on loss weig…

2024

DetKDS: Knowledge Distillation Search for Object Detectors

ICML 2024poster

In this paper, we present DetKDS, the first framework that searches for optimal detection distillation policies. Manual design of detection distillers becomes challenging and time-consuming due to significant disparities in distillation behaviors between detectors with different backbones, paradigms…

2024

Discovering Sparsity Allocation for Layer-wise Pruning of Large Language Models

NeurIPS 2024poster

In this paper, we present DSA, the first automated framework for discovering sparsity allocation schemes for layer-wise pruning in Large Language Models (LLMs). LLMs have become increasingly powerful, but their large parameter counts make them computationally expensive. Existing pruning methods fo…

Cited by 10SourcePDFScholar
2024

Era3D: High-Resolution Multiview Diffusion using Efficient Row-wise Attention

NeurIPS 2024poster

In this paper, we introduce **Era3D**, a novel multiview diffusion method that generates high-resolution multiview images from a single-view image. Despite significant advancements in multiview generation, existing methods still suffer from camera prior mismatch, inefficacy, and low resolution, resu…

Cited by 7SourcePDFScholar
2024

Weakly-Supervised Emotion Transition Learning for Diverse 3D Co-speech Gesture Generation

CVPR 2024poster

Generating vivid and emotional 3D co-speech gestures is crucial for virtual avatar animation in human-machine interaction applications. While the existing methods enable generating the gestures to follow a single emotion label they overlook that long gesture sequence modeling with emotion transition…

2023

F&F Attack: Adversarial Attack against Multiple Object Trackers by Inducing False Negatives and False Positives

ICCV 2023poster

Multi-object tracking (MOT) aims to build moving trajectories for number-agnostic objects. Modern multi-object trackers commonly follow the tracking-by-detection strategy. Therefore, fooling detectors can be an effective solution but it usually requires attacks in multiple successive frames, resulti…

Cited by 10PDFScholar
2023

Homography Guided Temporal Fusion for Road Line and Marking Segmentation

ICCV 2023poster

Reliable segmentation of road lines and markings is critical to autonomous driving. Our work is motivated by the observations that road lines and markings are (1) frequently occluded in the presence of moving vehicles, shadow, and glare and (2) highly structured with low intra-class shape variance a…

Cited by 5PDFcodeScholar
2023

InterTracker: Discovering and Tracking General Objects Interacting with Hands in the Wild

IROS 2023poster

Understanding human interaction with objects is an important research topic for embodied Artificial Intelligence and identifying the objects that humans are interacting with is a primary problem for interaction understanding. Existing methods rely on frame-based detectors to locate interacting objec…

Cited by 1SourceScholar
2023

MB-TaylorFormer: Multi-Branch Efficient Transformer Expanded by Taylor Formula for Image Dehazing

ICCV 2023poster

In recent years, Transformer networks are beginning to replace pure convolutional neural networks (CNNs) in the field of computer vision due to their global receptive field and adaptability to input. However, the quadratic computational complexity of softmax-attention limits the wide application in…

Cited by 124PDFcodeScholar
2023

PRIOR: Prototype Representation Joint Learning from Medical Images and Reports

ICCV 2023poster

Contrastive learning based vision-language joint pre-training has emerged as a successful representation learning strategy. In this paper, we present a prototype representation learning framework incorporating both global and local alignment between medical images and reports. In contrast to standar…

Cited by 63PDFcodeScholar
2023

Punctuation-level Attack: Single-shot and Single Punctuation Can Fool Text Models

NeurIPS 2023poster

The adversarial attacks have attracted increasing attention in various fields including natural language processing. The current textual attacking models primarily focus on fooling models by adding character-/word-/sentence-level perturbations, ignoring their influence on human perception. In this p…

Cited by 3SourcePDFScholar
2023

Robust Single Image Reflection Removal Against Adversarial Attacks

CVPR 2023poster

This paper addresses the problem of robust deep single-image reflection removal (SIRR) against adversarial attacks. Current deep learning based SIRR methods have shown significant performance degradation due to unnoticeable distortions and perturbations on input images. For a comprehensive robustnes…

2023

Ultra-High-Definition Low-Light Image Enhancement: A Benchmark and Transformer-Based Method

AAAI 2023technical

As the quality of optical sensors improves, there is a need for processing large-scale images. In particular, the ability of devices to capture ultra-high definition (UHD) images and video places new demands on the image processing pipeline. In this paper, we consider the task of low-light image enh…

2022

Aesthetic Text Logo Synthesis via Content-Aware Layout Inferring

CVPR 2022poster

Text logo design heavily relies on the creativity and expertise of professional designers, in which arranging element layouts is one of the most important procedures. However, few attention has been paid to this task which needs to take many factors (e.g., fonts, linguistics, topics, etc.) into cons…

Cited by 33PDFcodeScholar
2021

Benchmarking Ultra-High-Definition Image Super-Resolution

ICCV 2021poster

Increasingly, modern mobile devices allow capturing images at Ultra-High-Definition (UHD) resolution, which includes 4K and 8K images. However, current single image super-resolution (SISR) methods focus on super-resolving images to ones with resolution up to high definition (HD) and ignore higher-re…

Cited by 38PDFScholar
2021

Towards Distraction-Robust Active Visual Tracking

ICML 2021spotlight

In active visual tracking, it is notoriously difficult when distracting objects appear, as distractors often mislead the tracker by occluding the target or bringing a confusing appearance. To address this issue, we propose a mixed cooperative-competitive multi-agent game, where a target and multiple…

Cited by 45SourcePDFScholar
2020

Beyond Monocular Deraining: Stereo Image Deraining via Semantic Understanding

ECCV 2020poster

Rain is a common natural phenomenon. Taking images in the rain however often results in degraded quality of images, thus compromises the performance of many computer vision systems. Most existing de-rain algorithms use only one single input image and aim to recover a clean image. Few work has exploi…

Cited by 54SourcePDFScholar
2020

Fine-Grained Image-to-Image Transformation Towards Visual Recognition

CVPR 2020poster

Existing image-to-image transformation approaches primarily focus on synthesizing visually pleasing data. Generating images with correct identity labels is challenging yet much less explored. It is even more challenging to deal with image transformation tasks with large deformation in poses, viewpoi…

Cited by 35PDFScholar
2019

AD-VAT: An Asymmetric Dueling mechanism for learning Visual Active Tracking

ICLR 2019poster

Visual Active Tracking (VAT) aims at following a target object by autonomously controlling the motion system of a tracker given visual observations. Previous work has shown that the tracker can be trained in a simulator via reinforcement learning and deployed in real-world scenarios. However, during…

2019

Face Anti-Spoofing: Model Matters, so Does Data

CVPR 2019poster

Face anti-spoofing is an important task in full-stack face applications including face detection, verification, and recognition. Previous approaches build models on datasets which do not simulate the real-world data well (e.g., small scale, insignificant variance, etc.). Existing models may rely on…

Cited by 302PDFScholar
2019

Learning to Compose Dynamic Tree Structures for Visual Contexts

CVPR 2019oral

We propose to compose dynamic tree structures that place the objects in an image into a visual context, helping visual reasoning tasks such as scene graph generation and visual Q&A. Our visual context tree model, dubbed VCTree, has two key advantages over existing structured object representations i…

Cited by 618PDFScholar
2019

Liquid Warping GAN: A Unified Framework for Human Motion Imitation, Appearance Transfer and Novel View Synthesis

ICCV 2019poster

We tackle the human motion imitation, appearance transfer, and novel view synthesis within a unified framework, which means that the model once being trained can be used to handle all these tasks. The existing task-specific methods mainly use 2D keypoints (pose) to estimate the human body structure.…

Cited by 333PDFcodeScholar
2019

Residual Regression With Semantic Prior for Crowd Counting

CVPR 2019poster

Crowd counting is a challenging task due to factors such as large variations in crowdedness and severe occlusions. Although recent deep learning based counting algorithms have achieved a great progress, the correlation knowledge among samples and the semantic prior have not yet been fully exploited.…

Cited by 129PDFcodeScholar
2018

Bi-Real Net: Enhancing the Performance of 1-bit CNNs with Improved Representational Capability and Advanced Training Algorithm

ECCV 2018poster

In this work, we study the 1-bit convolutional neural networks (CNNs), of which both the weights and activations are binary. While being efficient, the classification accuracy of the current 1-bit CNNs is much worse compared with their counterpart real-valued CNN models on the large-scale dataset, l…

2018

End-to-end Active Object Tracking via Reinforcement Learning

ICML 2018oral

We study active object tracking, where a tracker takes as input the visual observation (i.e. frame sequence) and produces the camera control signal (e.g., move forward, turn left, etc). Conventional methods tackle the tracking and the camera control separately, which is challenging to tune jointly.…

Cited by 113SourcePDFScholar
2018

Learning to Generate Time-Lapse Videos Using Multi-Stage Dynamic Generative Adversarial Networks

CVPR 2018poster

Taking a photo outside, can we predict the immediate future, e.g., how would the cloud move in the sky? We address this problem by presenting a generative adversarial network (GAN) based two-stage approach to generating realistic time-lapse videos of high resolution. Given the first frame, our model…

Cited by 209SourcePDFScholar