← Search

Xiaoming Li

18 accepted papers

2026

From ``Sure" to ``Sorry": Detecting Jailbreak in Large Vision Language Model via JailNeurons

ICLR 2026poster

Large Vision-Language Models (LVLMs) are vulnerable to jailbreak attacks that can generate harmful content. Existing detection methods are either limited to detecting specific attack types or are too time-consuming, making them impractical for real-world deployment. To address these challenges, we p…

Cited by 0SourcecodeScholar
2026

RefSTAR: Blind Face Image Restoration with Reference Selection, Transfer, and Reconstruction

AAAI 2026technical

Introducing high-quality references can largely alleviate the uncertainty in blind face image restoration tasks, yet the equivocal utilization of reference priors makes it still a struggle to well preserve the human identity. We attribute the identity inconsistency to two deficiencies of existing re

Cited by 0SourcePDFScholar
2025

Omegance: A Single Parameter for Various Granularities in Diffusion-Based Synthesis

ICCV 2025poster

In this work, we show that we only need a single parameter \omega to effectively control granularity in diffusion-based synthesis. This parameter is incorporated during the denoising steps of the diffusion model's reverse process. This simple approach does not require model retraining or architectur…

2024

VQ-FONT: Few-Shot Font Generation with Structure-Aware Enhancement and Quantization

AAAI 2024technical

Few-shot font generation is challenging, as it needs to capture the fine-grained stroke styles from a limited set of reference glyphs, and then transfer to other characters, which are expected to have similar styles. However, due to the diversity and complexity of Chinese font styles, the synthesize…

2024

When StyleGAN Meets Stable Diffusion: a W+ Adapter for Personalized Image Generation

CVPR 2024poster

Text-to-image diffusion models have remarkably excelled in producing diverse high-quality and photo-realistic images. This advancement has spurred a growing interest in incorporating specific identities into generated content. Most current methods employ an inversion approach to embed a target visua…

2023

Learning Generative Structure Prior for Blind Text Image Super-Resolution

CVPR 2023poster

Blind text image super-resolution (SR) is challenging as one needs to cope with diverse font styles and unknown degradation. To address the problem, existing methods perform character recognition in parallel to regularize the SR task, either through a loss constraint or intermediate feature conditio…

2023

MetaF2N: Blind Image Super-Resolution by Learning Efficient Model Adaptation from Faces

ICCV 2023poster

Due to their highly structured characteristics, faces are easier to recover than natural scenes for blind image super-resolution. Therefore, we can extract the degradation representation of an image from the low-quality and recovered face pairs. Using the degradation representation, realistic low-qu…

Cited by 7PDFcodeScholar
2022

From Face to Natural Image: Learning Real Degradation for Blind Image Super-Resolution

ECCV 2022poster

"How to design proper training pairs is critical for super-resolving real-world low-quality (LQ) images, which suffers from the difficulties in either acquiring paired ground-truth high-quality (HQ) images or synthesizing photo-realistic degraded LQ observations. Recent works mainly focus on modelin…

2022

Semantic-Shape Adaptive Feature Modulation for Semantic Image Synthesis

CVPR 2022poster

Recent years have witnessed substantial progress in semantic image synthesis, it is still challenging in synthesizing photo-realistic images with rich details. Most previous methods focus on exploiting the given semantic map, which just captures an object-level layout for an image. Obviously, a fine…

Cited by 35PDFcodeScholar
2021

Learning Semantic Person Image Generation by Region-Adaptive Normalization

CVPR 2021poster

Human pose transfer has received great attention due to its wide applications, yet is still a challenging task that is not well solved. Recent works have achieved great success to transfer the person image from the source to the target pose. However, most of them cannot well capture the semantic app…

Cited by 81PDFcodeScholar
2021

Progressive Semantic-Aware Style Transformation for Blind Face Restoration

CVPR 2021poster

Face restoration is important in face image processing, and has been widely studied in recent years. However, previous works often fail to generate plausible high quality (HQ) results for real-world low quality (LQ) face images. In this paper, we propose a new progressive semantic-aware style transf…

Cited by 195PDFcodeScholar
2020

Blind Face Restoration via Deep Multi-scale Component Dictionaries

ECCV 2020poster

Recent reference-based face restoration methods have received considerable attention due to their great capability in recovering high-frequency details on real low-quality images. However, most of these methods require a high-quality reference image of the same identity, making them only applicable…

2020

Enhanced Blind Face Restoration With Multi-Exemplar Images and Adaptive Spatial Feature Fusion

CVPR 2020oral

In many real-world face restoration applications, e.g., smartphone photo albums and old films, multiple high-quality (HQ) images of the same person usually are available for a given degraded low-quality (LQ) observation. However, most existing guided face restoration methods are based on single HQ e…

Cited by 129PDFcodeScholar
2020

Face Super-Resolution Guided by 3D Facial Priors

ECCV 2020poster

State-of-the-art face super-resolution methods employ deep convolutional neural networks to learn a mapping between low- and high-resolution facial patterns by exploring local appearance knowledge. However, most of these methods do not well exploit facial structures and identity information, and str…

Cited by 85SourcePDFScholar
2020

Preference-Aware Mask for Session-Based Recommendation with Bidirectional Transformer

ICASSP 2020accepted

User profiles are not always visible in E-commerce scenarios, in which case the recommender systems can only summarize users' preferences through sessions of historical records. However, the items in a session might be irrelevant to users' preferences or become the disturbances for modelling the use…

Cited by 0SourceScholar
2018

Learning Warped Guidance for Blind Face Restoration

ECCV 2018poster

This paper studies the problem of blind face restoration from an unconstrained blurry, noisy, low-resolution, or compressed image (i.e., degraded observation). For better recovery of fine facial details, we modify the problem setting by taking both the degraded observation and a high-quality guided…

2018

Shift-Net: Image Inpainting via Deep Feature Rearrangement

ECCV 2018poster

Deep convolutional networks (CNNs) have exhibited their potential in image inpainting for producing plausible results. However, in most existing methods, e.g., context encoder, the missing parts are predicted by propagating the surrounding convolutional features through a fully connected layer, whic…