← Search

Wenbo Li

71 accepted papers

2026

FastFLUX: Pruning FLUX with Block-wise Replacement and Sandwich Training

AAAI 2026technical

Recent advancements in text-to-image (T2I) generation have led to the emergence of highly expressive models such as diffusion transformers (DiTs), exemplified by FLUX. However, their massive parameter sizes lead to slow inference, high memory usage, and poor deployability. Existing acceleration meth

Cited by 0SourcePDFScholar
2026

FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video Generation

AAAI 2026technical

DiT models have achieved great success in text-to-video generation, leveraging their scalability in model capacity and data scale. High content and motion fidelity aligned with text prompts, however, often require large model parameters and a substantial number of function evaluations (NFEs). Realis

Cited by 0SourcePDFScholar
2026

Grounding-IQA: Grounding Multimodal Language Model for Image Quality Assessment

ICLR 2026poster

The development of multimodal large language models (MLLMs) enables the evaluation of image quality through natural language descriptions. This advancement allows for more detailed assessments. However, these MLLM-based IQA methods primarily rely on general contextual descriptions, sometimes limitin…

Cited by 0SourcecodeScholar
2026

JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator Optimization

CVPR 2026

Agent-based editing models have substantially advanced interactive experiences, processing quality, and creative flexibility. However, two critical challenges persist: (1) instruction hallucination--text-only chain-of-thought (CoT) reasoning cannot fully prevent factual errors due to inherent inform

Cited by 0SourcecodeScholar
2026

Pixel to Gaussian: Ultra-Fast Continuous Super-Resolution with 2D Gaussian Modeling

ICLR 2026poster

Arbitrary-scale super-resolution (ASSR) aims to reconstruct high-resolution (HR) images from low-resolution (LR) inputs with arbitrary upsampling factors using a single model, addressing the limitations of traditional SR methods constrained to fixed-scale factors (\textit{e.g.}, $\times$ 2). Recent…

Cited by 0SourcecodeScholar
2026

QuantVSR: Low-Bit Post-Training Quantization for Real-World Video Super-Resolution

AAAI 2026technical

Diffusion models have shown superior performance in real-world video super-resolution (VSR). However, the slow processing speeds and heavy resource consumption of diffusion models hinder their practical application and deployment. Quantization offers a potential solution for compressing the VSR mode

Cited by 0SourcePDFScholar
2026

SODiff:Semantic-Oriented Diffusion Model for JPEG Compression Artifacts Removal

AAAI 2026technical

JPEG, as a widely used image compression standard, often introduces severe visual artifacts when achieving high compression ratios. Although existing deep learning-based restoration methods have made considerable progress, they often struggle to recover complex texture details, resulting in over-smo

Cited by 0SourcePDFScholar
2026

Test-Time Preference Optimization for Image Restoration

AAAI 2026technical

Image restoration (IR) models are typically trained to recover high-quality images using L1 or LPIPS loss. To handle diverse unknown degradations, zero-shot IR methods have also been introduced. However, existing pre-trained and zero-shot IR approaches often fail to align with human preferences, res

Cited by 0SourcePDFScholar
2026

UniRestorer: Universal Image Restoration via Adaptively Estimating Image Degradation at Proper Granularity

ICLR 2026poster

Recently, considerable progress has been made in all-in-one image restoration. Generally, existing methods can be degradation-agnostic or degradation-aware. However, the former are limited in leveraging degradation estimation-based priors, and the latter suffer from the inevitable error in degradati…

Cited by 0SourcecodeScholar
2025

AVG-LLaVA: An Efficient Large Multimodal Model with Adaptive Visual Granularity

ACL 2025finding

Recently, large multimodal models (LMMs) have achieved significant advancements. When dealing with high-resolution images, dominant LMMs typically divide them into multiple local images and a global image, leading to a large number of visual tokens. In this work, we introduce AVG-LLaVA, an LMM that…

2025

BiMaCoSR: Binary One-Step Diffusion Model Leveraging Flexible Matrix Compression for Real Super-Resolution

ICML 2025poster

While super-resolution (SR) methods based on diffusion models (DM) have demonstrated inspiring performance, their deployment is impeded due to the heavy request of memory and computation. Recent researchers apply two kinds of methods to compress or fasten the DM. One is to compress the DM into 1-bit…

2025

CamEdit: Continuous Camera Parameter Control for Photorealistic Image Editing

NeurIPS 2025poster

Recent advances in diffusion models have substantially improved text-driven image editing. However, existing frameworks based on discrete textual tokens struggle to support continuous control over camera parameters and smooth transitions in visual effects. These limitations hinder their applications…

Cited by 0SourceScholar
2025

Compression-Aware One-Step Diffusion Model for JPEG Artifact Removal

ICCV 2025poster

Diffusion models have demonstrated remarkable success in image restoration tasks. However, their multi-step denoising process introduces significant computational overhead, limiting their practical deployment. Furthermore, existing methods struggle to effectively remove severe JPEG artifact, especia…

2025

Directing Mamba to Complex Textures: An Efficient Texture-Aware State Space Model for Image Restoration

IJCAI 2025

Image restoration aims to recover details and enhance contrast in degraded images. With the growing demand for high-quality imaging (e.g., 4K and 8K), achieving a balance between restoration quality and computational efficiency has become increasingly critical. Existing methods, primarily based on C

Cited by 0SourcePDFScholar
2025

Dual Prompting Image Restoration with Diffusion Transformers

CVPR 2025poster

Recent state-of-the-art image restoration methods mostly adopt latent diffusion models with U-Net backbones, yet still facing challenges in achieving high-quality restoration due to their limited capabilities. Diffusion transformers (DiTs), like SD3, are emerging as a promising alternative because o…

Cited by 1SourcePDFScholar
2025

Dual-Mode Motion Control of Multi-Stimulus Deformable Miniature Robots with Adaptive Orientation Compensation in Unstructured Environments

IROS 2025

Miniature robots hold great promise for performing micromanipulation tasks within hard-to-reach confined spaces. However, effectively maneuvering across complex and unstructured terrain, achieving adaptive morphogenesis, and developing adaptive multimodal locomotion strategies remain challenges for

Cited by 0SourceScholar
2025

Fast Image Super-Resolution via Consistency Rectified Flow

ICCV 2025poster

Diffusion models (DMs) have demonstrated remarkable success in real-world image super-resolution (SR), yet their reliance on time-consuming multi-step sampling largely hinders their practical applications. While recent efforts have introduced few- or single-step solutions, existing methods either in…

Cited by 0SourcePDFScholar
2025

JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent

NeurIPS 2025poster

Photo retouching has become integral to contemporary visual storytelling, enabling users to capture aesthetics and express creativity. While professional tools such as Adobe Lightroom offer powerful capabilities, they demand substantial expertise and manual effort. In contrast, existing AI-based sol…

Cited by 0SourceScholar
2025

JarvisIR: Elevating Autonomous Driving Perception with Intelligent Image Restoration

CVPR 2025poster

Vision-centric perception systems often struggle with unpredictable and coupled weather degradations in the wild. Current solutions are often limited, as they either depend on specific degradation priors or suffer from significant domain gaps. To enable robust and autonomous operation in real-world…

Cited by 2SourcePDFScholar
2025

MC^2: Multi-concept Guidance for Customized Multi-concept Generation

CVPR 2025poster

Customized text-to-image generation, which synthesizes images based on user-specified concepts, has made significant progress in handling individual concepts. However, when extended to multiple concepts, existing methods often struggle with properly integrating different models and avoiding the unin…

2025

MambaIRv2: Attentive State Space Restoration

CVPR 2025poster

The Mamba-based image restoration backbones have recently demonstrated significant potential in balancing global reception and computational efficiency. However, the inherent causal modeling limitation of Mamba, where each token depends solely on its predecessors in the scanned sequence, restricts t…

2025

NopeRoomGS: Indoor 3D Gaussian Splatting Optimization without Camera Pose Input

NeurIPS 2025poster

Recent advances in 3D Gaussian Splatting (3DGS) have enabled real-time, high-fidelity view synthesis, but remain critically dependent on camera poses estimated by Structure-from-Motion (SfM), which is notoriously unreliable in textureless indoor environments. To eliminate this dependency, recent pos…

Cited by 0SourceScholar
2025

OSCAR: One-Step Diffusion Codec Across Multiple Bit-rates

NeurIPS 2025poster

Pretrained latent diffusion models have shown strong potential for lossy image compression, owing to their powerful generative priors. Most existing diffusion-based methods reconstruct images by iteratively denoising from random noise, guided by compressed latent representations. While these approac…

Cited by 0SourcecodeScholar
2025

One Diffusion Step to Real-World Super-Resolution via Flow Trajectory Distillation

ICML 2025poster

Diffusion models (DMs) have significantly advanced the development of real-world image super-resolution (Real-ISR), but the computational cost of multi-step diffusion models limits their application. One-step diffusion models generate high-quality images in a one sampling step, greatly reducing comp…

2025

PATIMT-Bench: A Multi-Scenario Benchmark for Position-Aware Text Image Machine Translation in Large Vision-Language Models

EMNLP 2025

Text Image Machine Translation (TIMT) aims to translate texts embedded within an image into another language. Current TIMT studies primarily focus on providing translations for all the text within an image, while neglecting to provide bounding boxes and covering limited scenarios. In this work, we e

2025

PMQ-VE: Progressive Multi-Frame Quantization for Video Enhancement

NeurIPS 2025poster

Multi-frame video enhancement tasks aim to improve the spatial and temporal resolution and quality of video sequences by leveraging temporal information from multiple frames, which are widely used in streaming video processing, surveillance, and generation. Although numerous Transformer-based enhanc…

Cited by 0SourcecodeScholar
2025

POSTA: A Go-to Framework for Customized Artistic Poster Generation

CVPR 2025poster

Poster design is a critical medium for visual communication. Prior work has explored automatic poster design using deep learning techniques, but these approaches lack text accuracy, user customization, and aesthetic appeal, limiting their applicability in artistic domains such as movies and exhibiti…

Cited by 4SourcePDFScholar
2025

PassionSR: Post-Training Quantization with Adaptive Scale in One-Step Diffusion based Image Super-Resolution

CVPR 2025poster

Diffusion-based image super-resolution (SR) models have shown superior performance at the cost of multiple denoising steps. However, even though the denoising step has been reduced to one, they require high computational costs and storage requirements, making it difficult for deployment on hardware…

2025

PocketSR: The Super-Resolution Expert in Your Pocket Mobiles

NeurIPS 2025poster

Real-world image super-resolution (RealSR) aims to enhance the visual quality of in-the-wild images, such as those captured by mobile phones. While existing methods leveraging large generative models demonstrate impressive results, the high computational cost and latency make them impractical for ed…

Cited by 0SourceScholar
2025

QMambaBSR: Burst Image Super-Resolution with Query State Space Model

CVPR 2025poster

Burst super-resolution (BurstSR) aims to reconstruct high-resolution images by fusing subpixel details from multiple low-resolution burst frames. The primary challenge lies in effectively extracting useful information while mitigating the impact of high-frequency noise. Most existing methods rely on…

Cited by 5SourcePDFScholar
2025

ResMaster: Mastering High-Resolution Image Generation via Structural and Fine-Grained Guidance

AAAI 2025technical

Diffusion models excel at producing high-quality images; however, scaling to higher resolutions, such as 4K, often results in structural distortions, and repetitive patterns. To this end, we introduce ResMaster, a novel, training-free method that empowers resolution-limited diffusion models to gener…

Cited by 11SourcePDFScholar
2025

Restabilizing Diffusion Models with Predictive Noise Fusion Strategy for Image Super-Resolution

AAAI 2025technical

Diffusion models are prominent in image generation for producing detailed and realistic images from Gaussian noises. However, they often encounter instability issues in image restoration tasks, e.g., super-resolution. Existing methods typically rely on multiple runs to find an initial noise that pro…

2025

Segment Any-Quality Images with Generative Latent Space Enhancement

CVPR 2025poster

Despite their success, Segment Anything Models (SAMs) experience significant performance drops on severely degraded, low-quality images, limiting their effectiveness in real-world scenarios. To address this, we propose GleSAM, which utilizes Generative Latent space Enhancement to boost robustness on…

Cited by 0SourcePDFScholar
2025

Towards Realistic Data Generation for Real-World Super-Resolution

ICLR 2025poster

Existing image super-resolution (SR) techniques often fail to generalize effectively in complex real-world settings due to the significant divergence between training data and practical scenarios. To address this challenge, previous efforts have either manually simulated intricate physical-based deg…

Cited by 14SourcePDFScholar
2025

Turbo2K: Towards Ultra-Efficient and High-Quality 2K Video Synthesis

ICCV 2025poster

Demand for 2K video synthesis is rising with increasing consumer expectations for ultra-clear visuals.While diffusion transformers (DiTs) have demonstrated remarkable capabilities in high-quality video generation, scaling them to 2K resolution remains computationally prohibitive due to quadratic gro…

Cited by 0SourcePDFScholar
2025

TurboVSR: Fantastic Video Upscalers and Where to Find Them

ICCV 2025poster

Diffusion-based generative models have demonstrated exceptional promise in the video super-resolution (VSR) task, achieving a substantial advancement in detail generation relative to prior methods. However, these approaches face significant computational efficiency challenges. For instance, current…

Cited by 0SourcePDFScholar
2025

Unsupervised Diffusion-Based Degradation Modeling for Real-World Super-Resolution

AAAI 2025technical

Single image super-solution (SR) aims to restore a high-resolution (HR) image from a degraded low-resolution (LR) image. However, existing SR models still face a significant domain gap between synthetic and real-world datasets due to the mismatched degradation distributions, hindering SR models from…

2024

A Manta Ray-Inspired Fast-Swimming Soft Electrohydraulic Robotic Fish

RA-L 2024

Underwater soft robots inspired by marine life have shown great potential in ocean exploration, monitoring, scientific research, etc., due to their excellent safety, compatibility and adaptability when interacting with underwater environments. However, most of their soft actuators suffer performance

Cited by 11SourceScholar
2024

Adaptive Meta-Learning Probabilistic Inference Framework for Long Sequence Prediction

AAAI 2024technical

Long sequence prediction has broad and significant application value in fields such as finance, wind power, and weather. However, the complex long-term dependencies of long sequence data and the potential domain shift problems limit the effectiveness of traditional models in practical scenarios. To…

2024

CoSeR: Bridging Image and Language for Cognitive Super-Resolution

CVPR 2024poster

Existing super-resolution (SR) models primarily focus on restoring local texture details often neglecting the global semantic information within the scene. This oversight can lead to the omission of crucial semantic details or the introduction of inaccurate textures during the recovery process. In o…

2024

Empowering Backbone Models for Visual Text Generation with Input Granularity Control and Glyph-Aware Training

EMNLP 2024main

Diffusion-based text-to-image models have demonstrated impressive achievements in diversity and aesthetics but struggle to generate images with legible visual texts. Existing backbone models have limitations such as misspelling, failing to generate texts, and lack of support for Chinese texts, but t…

2024

Image Inpainting via Iteratively Decoupled Probabilistic Modeling

ICLR 2024spotlight

Generative adversarial networks (GANs) have made great success in image inpainting yet still have difficulties tackling large missing regions. In contrast, iterative probabilistic algorithms, such as autoregressive and denoising diffusion models, have to be deployed with massive computing resources…

Cited by 11SourcePDFScholar
2024

Low-Res Leads the Way: Improving Generalization for Super-Resolution by Self-Supervised Learning

CVPR 2024poster

For image super-resolution (SR) bridging the gap between the performance on synthetic datasets and real-world degradation scenarios remains a challenge. This work introduces a novel "Low-Res Leads the Way" (LWay) training framework merging Supervised Pre-training with Self-supervised Learning to enh…

Cited by 15SourcePDFScholar
2024

RestoreAgent: Autonomous Image Restoration Agent via Multimodal Large Language Models

NeurIPS 2024poster

Natural images captured by mobile devices often suffer from multiple types of degradation, such as noise, blur, and low light. Traditional image restoration methods require manual selection of specific tasks, algorithms, and execution sequences, which is time-consuming and may yield suboptimal resul…

Cited by 6SourcePDFScholar
2024

UltraPixel: Advancing Ultra High-Resolution Image Synthesis to New Peaks

NeurIPS 2024poster

Ultra-high-resolution image generation poses great challenges, such as increased semantic planning complexity and detail synthesis difficulties, alongside substantial training resource demands. We present UltraPixel, a novel architecture utilizing cascade diffusion models to generate high-quality im…

Cited by 17SourcePDFScholar
2024

Unleashing Multispectral Video's Potential in Semantic Segmentation: A Semi-supervised Viewpoint and New UAV-View Benchmark

NeurIPS 2024poster

Thanks to the rapid progress in RGB & thermal imaging, also known as multispectral imaging, the task of multispectral video semantic segmentation, or MVSS in short, has recently drawn significant attentions. Noticeably, it offers new opportunities in improving segmentation performance under unfavora…

Cited by 0SourcePDFScholar
2024

Unveiling Advanced Frequency Disentanglement Paradigm for Low-Light Image Enhancement

ECCV 2024poster

"Previous low-light image enhancement (LLIE) approaches, while employing frequency decomposition techniques to address the intertwined challenges of low frequency (e.g., illumination recovery) and high frequency (e.g., noise reduction), primarily focused on the development of dedicated and complex n…

2023

Exploring Motion Ambiguity and Alignment for High-Quality Video Frame Interpolation

CVPR 2023poster

For video frame interpolation(VFI), existing deep-learning-based approaches strongly rely on the ground-truth (GT) intermediate frames, which sometimes ignore the non-unique nature of motion judging from the given adjacent frames. As a result, these methods tend to produce averaged solutions that ar…

Cited by 26SourcePDFScholar
2023

NeRFLix: High-Quality Neural View Synthesis by Learning a Degradation-Driven Inter-Viewpoint MiXer

CVPR 2023poster

Neural radiance fields(NeRF) show great success in novel-view synthesis. However, in real-world scenes, recovering high-quality details from the source images is still challenging for the existing NeRF-based approaches, due to the potential imperfect calibration information and scene representation…

2023

On Efficient Transformer-Based Image Pre-training for Low-Level Vision

IJCAI 2023poster

Pre-training has marked numerous state of the arts in high-level computer vision, while few attempts have ever been made to investigate how pre-training acts in image processing systems. In this paper, we tailor transformer-based pre-training regimes that boost various low-level tasks. To comprehens…

2023

Texture Generation on 3D Meshes with Point-UV Diffusion

ICCV 2023oral

In this work, we focus on synthesizing high-quality textures on 3D meshes. We present Point-UV diffusion, a coarse-to-fine pipeline that marries the denoising diffusion model with UV mapping to generate 3D consistent and high-quality texture images in UV space. We start with introducing a point diff…

Cited by 45PDFcodeScholar
2022

Best-Buddy GANs for Highly Detailed Image Super-resolution

AAAI 2022technical

We consider the single image super-resolution (SISR) problem, where a high-resolution (HR) image is generated based on a low-resolution (LR) input. Recently, generative adversarial networks (GANs) become popular to hallucinate details. Most methods along this line rely on a predefined single-LR-sing…

2022

SceneSqueezer: Learning To Compress Scene for Camera Relocalization

CVPR 2022oral

Standard visual localization methods build a priori 3D model of a scene which is used to establish correspondences against the 2D keypoints in a query image. Storing these pre-built 3D scene models can be prohibitively expensive for large-scale environments, especially on mobile devices with limited…

Cited by 37PDFScholar
2022

Towards Efficient and Scale-Robust Ultra-High-Definition Image Demoiréing

ECCV 2022poster

"With the rapid development of mobile devices, modern widely-used mobile phones typically allow users to capture 4K resolution (i.e., ultra-high-definition) images. However, for image demoiréing, a challenging task in low-level vision, existing works are generally carried out on low-resolution or sy…

2022

Video Demoireing With Relation-Based Temporal Consistency

CVPR 2022poster

Moire patterns, appearing as color distortions, severely degrade the image and video qualities when filming a screen with digital cameras. Considering the increasing demands for capturing videos, we study how to remove such undesirable moire patterns in videos, namely video demoireing. To this end,…

Cited by 27PDFcodeScholar
2021

MASA-SR: Matching Acceleration and Spatial Adaptation for Reference-Based Image Super-Resolution

CVPR 2021poster

Reference-based image super-resolution (RefSR) has shown promising success in recovering high-frequency details by utilizing an external reference image (Ref). In this task, texture details are transferred from the Ref image to the low-resolution (LR) image according to their point- or patch-wise co…

Cited by 178PDFcodeScholar
2020

Conditional Image Repainting via Semantic Bridge and Piecewise Value Function

ECCV 2020poster

We study conditional image repainting where a model is trained to generate visual content conditioned on user inputs, and composite the generated content seamlessly onto a user provided image while preserving the semantics of users' inputs. The content generation community have been pursuing to lowe…

Cited by 6SourcePDFScholar
2020

Explainable and Efficient Sequential Correlation Network for 3D Single Person Concurrent Activity Detection

IROS 2020poster

We present the sequential correlation network (SCN) to improve concurrent activity detection. SCN combines a recurrent neural network and a correlation model hierarchically to model the complex correlations and temporal dynamics of concurrent activities. SCN has several advantages that enable effect…

Cited by 2SourceScholar
2020

LAPAR: Linearly-Assembled Pixel-Adaptive Regression Network for Single Image Super-resolution and Beyond

NeurIPS 2020poster

Single image super-resolution (SISR) deals with a fundamental problem of upsampling a low-resolution (LR) image to its high-resolution (HR) version. Last few years have witnessed impressive progress propelled by deep learning methods. However, one critical challenge faced by existing methods is to s…

2020

MISC: Multi-Condition Injection and Spatially-Adaptive Compositing for Conditional Person Image Synthesis

CVPR 2020poster

In this paper, we explore synthesizing person images with multiple conditions for various backgrounds. To this end, we propose a framework named "MISC" for conditional image generation and image compositing. For conditional image generation, we improve the existing condition injection mechanisms by…

Cited by 37PDFScholar
2020

MuCAN: Multi-Correspondence Aggregation Network for Video Super-Resolution

ECCV 2020poster

Video super-resolution (VSR) aims to utilize multiple low-resolution frames to generate a high-resolution prediction for each frame. In this process, inter- and intra-frames are the key sources for exploiting temporal and spatial information. However, there are a couple of limitations for existing V…

2019

Object-Driven Text-To-Image Synthesis via Adversarial Training

CVPR 2019poster

In this paper, we propose Object-driven Attentive Generative Adversarial Newtorks (Obj-GANs) that allow attention-driven, multi-stage refinement for synthesizing complex images from text descriptions. With a novel object-driven attentive generative network, the Obj-GAN can synthesize salient objects…

Cited by 386PDFScholar
2017

Adaptive RNN Tree for Large-Scale Human Action Recognition

ICCV 2017poster

In this work, we present the RNN Tree (RNN-T), an adaptive learning framework for skeleton based human action recognition. Our method categorizes action classes and uses multiple Recurrent Neural Networks (RNNs) in a tree-like hierarchy. The RNNs in RNN-T are co-trained with the action category hier…

Cited by 140PDFScholar