← Search

Xiaogang Xu

51 accepted papers

2026

Class Incremental Medical Image Segmentation via Prototype-Guided Calibration and Dual-Aligned Distillation

AAAI 2026technical

Class incremental medical image segmentation (CIMIS) aims to preserve knowledge of previously learned classes while learning new ones without relying on old-class annotations. However, existing methods 1) either adopt one-size-fits-all strategies that treat all spatial regions and feature channels e

Cited by 0SourcePDFScholar
2026

FC-VFI: FAITHFUL AND CONSISTENT VIDEO FRAME INTERPOLATION FOR HIGH-FPS SLOW MOTION VIDEO GENERATION

ICASSP 2026oral

Large pre-trained video diffusion models excel in video frame interpolation but struggle to generate high fidelity frames due to reliance on intrinsic generative priors, limiting detail preservation from start and end frames. Existing methods often depend on motion control for temporal consistency,…

Cited by 0SourcePDFScholar
2026

Fair in Mind, Fair in Action? A Synchronous Benchmark for Understanding and Generation in UMLLMs

ICLR 2026poster

As artificial intelligence (AI) is increasingly deployed across domains, ensuring fairness has become a core challenge. However, the field faces a "Tower of Babel'' dilemma: fairness metrics abound, yet their underlying philosophical assumptions often conflict, hindering unified paradigms—particular…

Cited by 0SourceScholar
2026

HoloFair: Unified T2I Fairness Evaluation and Fair-GRPO Debiasing

ICML 2026poster

Text-to-Image (T2I) models have made significant strides in visual realism and semantic consistency, yet they often perpetuate and amplify societal biases. Existing evaluation methods typically address only single-dimensional biases, lacking perspectives to uncover model biases at social-related dee…

Cited by 0SourceScholar
2026

Learning Latent Proxies for Controllable Single-Image Relighting

CVPR 2026

Single-image relighting is highly under-constrained: small illumination changes can produce large, nonlinear variations in shading, shadows, and specularities, while geometry and materials remain unobserved. Existing diffusion-based approaches either rely on intrinsic- or G-buffer-based pipelines th

Cited by 0SourceScholar
2026

Robust-R1: Degradation-Aware Reasoning for Robust Visual Understanding

AAAI 2026technical

Multimodal Large Language Models struggle to maintain reliable performance under extreme real-world visual degradations, which impede their practical robustness. Existing robust MLLMs predominantly rely on implicit training/adaptation that focuses solely on visual encoder generalization, suffering f

Cited by 0SourcePDFScholar
2025

Co-Painter: Fine-Grained Controllable Image Stylization via Implicit Decoupling and Adaptive Injection

ICCV 2025poster

Controllable diffusion models have been widely applied in image stylization. However, existing methods often treat the style in the reference image as a single, indivisible entity, which makes it difficult to transfer specific stylistic attributes. To address this issue, we propose a fine-grained co…

2025

DR-Encoder: Encode Low-rank Gradients with Random Prior for Large Language Models Differentially Privately

AAAI 2025technical

The emergence of the large language model (LLM) has shown its superiority in a wide range of disciplines, including language understanding and translation, relational logic reasoning, and even partial differential equations solving. The transformer is the pervasive backbone architecture for the foun…

2025

DiMSOD: A Diffusion-Based Framework for Multi-Modal Salient Object Detection

AAAI 2025technical

Multi-modal salient object detection (SOD) through the integration of additional data such as depth or thermal information has become a significant task in computer vision during recent years. Traditionally, the challenges of identifying salient objects in RGB, RGB-D (Depth), and RGB-T (Thermal) ima…

Cited by 0SourcePDFScholar
2025

DiffDoctor: Diagnosing Image Diffusion Models Before Treating

ICCV 2025poster

In spite of recent progress, image diffusion models still produce artifacts. A common solution is to leverage the feedback provided by quality assessment systems or human annotators to optimize the model, where images are generally rated in their entirety. In this work, we believe problem-solving st…

Cited by 0SourcePDFScholar
2025

LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence

ICML 2025poster

Recent embodied agents are primarily built based on reinforcement learning (RL) or large language models (LLMs). Among them, RL agents are efficient for deployment but only perform very few tasks. By contrast, giant LLM agents (often more than 1000B parameters) present strong generalization while de…

Cited by 1SourcePDFScholar
2025

Learnable Feature Patches and Vectors for Boosting Low-light Image Enhancement without External Knowledge

ICCV 2025poster

A major challenge in Low-Light Image Enhancement (LLIE) is its ill-posed nature: low-light images often lack sufficient information to align with normal-light ones (e.g., not all training data can be fully fitted to the ground truth). Numerous studies have attempted to bridge the gap between low- an…

Cited by 0SourcePDFScholar
2025

Low-Light Video Enhancement via Spatial-Temporal Consistent Decomposition

IJCAI 2025

Low-Light Video Enhancement (LLVE) seeks to restore dynamic or static scenes plagued by severe invisibility and noise. In this paper, we present an innovative video decomposition strategy that incorporates view-independent and view-dependent components to enhance the performance of LLVE. We leverage

Cited by 0SourcePDFScholar
2025

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning

NeurIPS 2025poster

This work explores enabling Chain-of-Thought (CoT) reasoning to link visual cues across multiple images. A straightforward solution is to adapt rule-based reinforcement learning for Vision-Language Models (VLMs). However, such methods typically rely on manually curated question-answer pairs, which c…

Cited by 0SourceScholar
2025

Towards Understanding the Robustness of Diffusion-Based Purification: A Stochastic Perspective

ICLR 2025poster

Diffusion-Based Purification (DBP) has emerged as an effective defense mechanism against adversarial attacks. The success of DBP is often attributed to the forward diffusion process, which reduces the distribution gap between clean and adversarial images by adding Gaussian noise. Although this expla…

Cited by 0SourcePDFScholar
2025

Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance

NeurIPS 2025poster

We present Wan-Move, a simple and scalable framework that brings motion control to video generative models. Existing motion-controllable methods typically suffer from coarse control granularity and limited scalability, leaving their outputs insufficient for practical use. We narrow this gap by achie…

Cited by 0SourceScholar
2024

"Refine, Discriminate and Align: Stealing Encoders via Sample-Wise Prototypes and Multi-Relational Extraction"

ECCV 2024poster

"This paper introduces RDA, a pioneering approach designed to address two primary deficiencies prevalent in previous endeavors aiming at stealing pre-trained encoders: (1) suboptimal performances attributed to biased optimization objectives, and (2) elevated query costs stemming from the end-to-end…

2024

An Incremental Unified Framework for Small Defect Inspection

ECCV 2024poster

"Artificial Intelligence (AI)-driven defect inspection is pivotal in industrial manufacturing. However, existing inspection systems are typically designed for specific industrial products and struggle with diverse product portfolios and evolving processes. Although some previous studies attempt to a…

2024

Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data

CVPR 2024poster

This work presents Depth Anything a highly practical solution for robust monocular depth estimation. Without pursuing novel technical modules we aim to build a simple yet powerful foundation model dealing with any images under any circumstances. To this end we scale up the dataset by designing a dat…

2024

Generative Active Learning for Long-tailed Instance Segmentation

ICML 2024poster

Recently, large-scale language-image generative models have gained widespread attention and many works have utilized generated data from these models to further enhance the performance of perception tasks. However, not all generated data can positively impact downstream models, and these methods do…

2024

HAWK: Learning to Understand Open-World Video Anomalies

NeurIPS 2024poster

Video Anomaly Detection (VAD) systems can autonomously monitor and identify disturbances, reducing the need for manual labor and associated costs. However, current VAD systems are often limited by their superficial semantic understanding of scenes and minimal user interaction. Additionally, the prev…

2024

HELPD: Mitigating Hallucination of LVLMs by Hierarchical Feedback Learning with Vision-enhanced Penalty Decoding

EMNLP 2024main

Large Vision-Language Models (LVLMs) have shown remarkable performance on many visual-language tasks. However, these models still suffer from multimodal hallucination, which means the generation of objects or content that violates the images. Many existing work detects hallucination by directly judg…

2024

Learning to Remove Wrinkled Transparent Film with Polarized Prior

CVPR 2024poster

In this paper we study a new problem Film Removal (FR) which attempts to remove the interference of wrinkled transparent films and reconstruct the original information under films for industrial recognition systems. We first physically model the imaging of industrial materials covered by the film. C…

2024

LucidDreamer: Towards High-Fidelity Text-to-3D Generation via Interval Score Matching

CVPR 2024highlight

The recent advancements in text-to-3D generation mark a significant milestone in generative models unlocking new possibilities for creating imaginative 3D assets across various real-world scenarios. While recent advancements in text-to-3D generation have shown promise they often fall short in render…

2024

S2WAT: Image Style Transfer via Hierarchical Vision Transformer Using Strips Window Attention

AAAI 2024technical

Transformer's recent integration into style transfer leverages its proficiency in establishing long-range dependencies, albeit at the expense of attenuated local modeling. This paper introduces Strips Window Attention Transformer (S2WAT), a novel hierarchical vision transformer designed for style tr…

2024

Unveiling Advanced Frequency Disentanglement Paradigm for Low-Light Image Enhancement

ECCV 2024poster

"Previous low-light image enhancement (LLIE) approaches, while employing frequency decomposition techniques to address the intertwined challenges of low frequency (e.g., illumination recovery) and high frequency (e.g., noise reduction), primarily focused on the development of dedicated and complex n…

2023

CorresNeRF: Image Correspondence Priors for Neural Radiance Fields

NeurIPS 2023poster

Neural Radiance Fields (NeRFs) have achieved impressive results in novel view synthesis and surface reconstruction tasks. However, their performance suffers under challenging scenarios with sparse input views. We present CorresNeRF, a novel method that leverages image correspondence priors computed…

2023

Deep Parametric 3D Filters for Joint Video Denoising and Illumination Enhancement in Video Super Resolution

AAAI 2023technical

Despite the quality improvement brought by the recent methods, video super-resolution (SR) is still very challenging, especially for videos that are low-light and noisy. The current best solution is to subsequently employ best models of video SR, denoising, and illumination enhancement, but doing so…

2023

FreeMask: Synthetic Images with Dense Annotations Make Stronger Segmentation Models

NeurIPS 2023poster

Semantic segmentation has witnessed tremendous progress due to the proposal of various advanced network architectures. However, they are extremely hungry for delicate annotations to train, and the acquisition is laborious and unaffordable. Therefore, we present FreeMask in this work, which resorts t…

2023

Input-Dependent Dynamical Channel Association For Knowledge Distillation

ICASSP 2023accepted

Feature-map based knowledge distillation has exhibited its significance in improving the performance of student model. Existing works mainly focus on the formulation of knowledge, but ignore the number difference of channels due to heterogeneous architectures of teacher-student pair. They generally…

Cited by 0SourceScholar
2023

Out-of-Domain GAN Inversion via Invertibility Decomposition for Photo-Realistic Human Face Manipulation

ICCV 2023poster

The fidelity of Generative Adversarial Networks (GAN) inversion is impeded by Out-Of-Domain (OOD) areas (e.g., background, accessories) in the image. Detecting the OOD areas beyond the generation ability of the pre-trained model and blending these regions with the input image can enhance fidelity.…

Cited by 4PDFcodeScholar
2023

Point2Pix: Photo-Realistic Point Cloud Rendering via Neural Radiance Fields

CVPR 2023poster

Synthesizing photo-realistic images from a point cloud is challenging because of the sparsity of point cloud representation. Recent Neural Radiance Fields and extensions are proposed to synthesize realistic images from 2D input. In this paper, we present Point2Pix as a novel point renderer to link t…

Cited by 20SourcePDFScholar
2022

DecoupleNet: Decoupled Network for Domain Adaptive Semantic Segmentation

ECCV 2022poster

"Unsupervised domain adaptation in semantic segmentation alleviates the reliance on expensive pixel-wise annotation. It uses a labeled source domain dataset as well as unlabeled target domain images to learn a segmentation network. In this paper, we observe two main issues of existing domain-invaria…

2022

MTFormer: Multi-task Learning via Transformer and Cross-Task Reasoning

ECCV 2022poster

"In this paper, we explore the advantages of utilizing transformer structures for addressing multi-task learning (MTL). Specifically, we demonstrate that models with transformer structures are more appropriate for MTL than convolutional neural networks (CNNs), and we propose a novel transformer-base…

Cited by 67SourcePDFScholar
2021

Dynamic Divide-and-Conquer Adversarial Training for Robust Semantic Segmentation

ICCV 2021poster

Adversarial training is promising for improving robustness of deep neural networks towards adversarial perturbations, especially on the classification task. The effect of this type of training on semantic segmentation, contrarily, just commences. We make the initial attempt to explore the defense st…

Cited by 47PDFcodeScholar
2021

Seeing Dynamic Scene in the Dark: A High-Quality Video Dataset With Mechatronic Alignment

ICCV 2021poster

Low-light video enhancement is an important task. Previous work is mostly trained on paired static images or videos. We compile a new dataset formed by our new strategy that contains high-quality spatially-aligned video pairs from dynamic scenes in low- and normal-light conditions. We built it using…

Cited by 117PDFcodeScholar
2019

Homomorphic Latent Space Interpolation for Unpaired Image-To-Image Translation

CVPR 2019oral

Generative adversarial networks have achieved great success in unpaired image-to-image translation. Cycle consistency allows modeling the relationship between two distinct domains without paired data. In this paper, we propose an alternative framework, as an extension of latent space interpolation,…

Cited by 81PDFScholar