← Search

Xiangtao Kong

11 accepted papers

2026

VOSR: A Vision-Only Generative Model for Image Super-Resolution

CVPR 2026

Large-scale pre-trained text-to-image (T2I) diffusion models, such as Stable Diffusion, can be finetuned for image super-resolution (SR) with highly realistic details. While impressive, pre-training such multi-modal models demands billions of high-quality text-image pairs and substantial computation

Cited by 0SourcecodeScholar
2025

InstructRestore: Region-Customized Image Restoration with Human Instructions

NeurIPS 2025poster

Despite the significant progress in diffusion prior-based image restoration for real-world scenarios, most existing methods apply uniform processing to the entire image, lacking the capability to perform region-customized image restoration according to user preferences. In this work, we propose a ne…

Cited by 0SourcecodeScholar
2025

Toward Generalizing Visual Brain Decoding to Unseen Subjects

ICLR 2025poster

Visual brain decoding aims to decode visual information from human brain activities. Despite the great progress, one critical limitation of current brain decoding research lies in the lack of generalization capability to unseen subjects. Prior work typically focuses on decoding brain activity of ind…

2025

Training on the Benchmark Is Not All You Need

AAAI 2025technical

The success of Large Language Models (LLMs) relies heavily on the huge amount of pre-training data learned in the pre-training phase. The opacity of the pre-training process and the training data causes the results of many benchmark tests to become unreliable. If any model has been trained on a benc…

2024

E-EVAL: A Comprehensive Chinese K-12 Education Evaluation Benchmark for Large Language Models

ACL 2024findings

The rapid development of Large Language Models (LLMs) has led to their increasing utilization in Chinese K-12 education. Despite the growing integration of LLMs and education, the absence of a dedicated benchmark for evaluating LLMs within this domain presents a pressing concern. Consequently, there…

2024

Scaling Up to Excellence: Practicing Model Scaling for Photo-Realistic Image Restoration In the Wild

CVPR 2024poster

We introduce SUPIR (Scaling-UP Image Restoration) a groundbreaking image restoration method that harnesses generative prior and the power of model scaling up. Leveraging multi-modal techniques and advanced generative prior SUPIR marks a significant advance in intelligent and realistic image restorat…

Cited by 49SourcePDFScholar
2023

DegAE: A New Pretraining Paradigm for Low-Level Vision

CVPR 2023highlight

Self-supervised pretraining has achieved remarkable success in high-level vision, but its application in low-level vision remains ambiguous and not well-established. What is the primitive intention of pretraining? What is the core problem of pretraining in low-level vision? In this paper, we aim to…

2023

Networks are Slacking Off: Understanding Generalization Problem in Image Deraining

NeurIPS 2023poster

Deep deraining networks consistently encounter substantial generalization issues when deployed in real-world applications, although they are successful in laboratory benchmarks. A prevailing perspective in deep learning encourages using highly complex data for training, with the expectation that ric…

Cited by 8SourcePDFScholar
2023

Retiformer: Retinex-Based Enhancement In Transformer For Low-Light Image

ICASSP 2023accepted

Transformer-based methods have shown impressive potential in many low-level vision tasks but are rarely used for low-light image enhancement (LLIE). Direct use of Transformer in LLIE will bring unnatural visual effects. This phenomenon encourages us to attempt to learn from the theory of Retinex. Af…

Cited by 0SourceScholar
2021

ClassSR: A General Framework to Accelerate Super-Resolution Networks by Data Characteristic

CVPR 2021poster

We aim at accelerating super-resolution (SR) networks on large images (2K-8K). The large images are usually decomposed into small sub-images in practical usages. Based on this processing, we found that different image regions have different restoration difficulties and can be processed by networks w…

Cited by 217PDFcodeScholar