← Search

Chunxia Xiao

26 accepted papers

2026

HumanPro: Single-view 3D Clothed Human Reconstruction with Progressive Normal Guidance

AAAI 2026technical

Reconstructing fine-grained geometry of clothed human from single-view image is a challenging task, particularly in accurately recovering complex shapes and generating clothes details. To address these limitations, we propose a novel approach named HumanPro, which estimates high-quality human normal

Cited by 0SourcePDFScholar
2026

Spherical Physics-Informed Neural Operator with Multi-Scale Coupling for Meteorological Downscaling

IJCAI 2026

Meteorological downscaling is crucial for high-resolution regional climate forecasting and disaster early warning. While neural operators have emerged as a promising paradigm for modeling complex spatiotemporal mappings, existing frameworks often struggle with spherical manifold geometric distortion

Cited by 0Scholar
2025

FriendsQA: A New Large-Scale Deep Video Understanding Dataset with Fine-grained Topic Categorization for Story Videos

AAAI 2025technical

Video question answering (VideoQA) aims to answer natural language questions according to the given videos. Although existing models perform well in the factoid VideoQA task, they still face challenges in deep video understanding (DVU) task, which focuses on story videos. Compared to factoid videos,…

2025

GGS: Generalizable Gaussian Splatting for Lane Switching in Autonomous Driving

AAAI 2025technical

We propose GGS, a Generalizable Gaussian Splatting method for Autonomous Driving that can achieve realistic rendering under large viewpoint changes. Previous generalizable 3D gaussian splatting methods are limited to rendering novel views that are very close to the original pair of images, which can…

Cited by 1SourcePDFScholar
2025

Hierarchical Adaptive Filtering Network for Text Image Specular Highlight Removal

CVPR 2025poster

Despite significant advances in the field of specular highlight removal in recent years, existing methods predominantly focus on natural images, where highlights typically appear on raised or edged surfaces of objects. These highlights are often small and sparsely distributed. However, for text imag…

Cited by 0SourcePDFScholar
2025

PHR-DIFF: Portrait Highlights Removal via Patch-aware Diffusion Model

AAAI 2025technical

Portraits often suffer from specular highlights due to factors like skin oiliness, lighting conditions, and shooting angles, which degrade aesthetics and affect downstream tasks. Thus, portrait highlight removal is imperative. Previous methods struggle to remove highlights and achieve high-fidelity…

Cited by 0SourcePDFScholar
2025

Rethinking the Adversarial Robustness of Multi-Exit Neural Networks in an Attack-Defense Game

CVPR 2025poster

Multi-exit neural networks represent a promising approach to enhancing model inference efficiency, yet like common neural networks, they suffer from significantly reduced robustness against adversarial attacks. While some defense methods have been raised to strengthen the adversarial robustness of m…

Cited by 0SourcePDFScholar
2025

iG-6DoF: Model-free 6DoF Pose Estimation for Unseen Object via Iterative 3D Gaussian Splatting

CVPR 2025poster

Traditional methods in pose estimation often rely on precise 3D models or additional data such as depth and normals, limiting their generalization, especially when objects undergo large translations or rotations. We propose iG-6DoF, a novel model-free 6D pose estimation method iterative 3D Gaussian…

Cited by 0SourcePDFScholar
2024

DLCA-Recon: Dynamic Loose Clothing Avatar Reconstruction from Monocular Videos

AAAI 2024technical

Reconstructing a dynamic human with loose clothing is an important but difficult task. To address this challenge, we propose a method named DLCA-Recon to create human avatars from monocular videos. The distance from loose clothing to the underlying body rapidly changes in every frame when the human…

Cited by 2SourcePDFScholar
2024

Diffusion-FOF: Single-View Clothed Human Reconstruction via Diffusion-Based Fourier Occupancy Field

CVPR 2024poster

Reconstructing a clothed human from a single-view image has several challenging issues including flexibly representing various body shapes and poses estimating complete 3D geometry and consistent texture and achieving more fine-grained details. To address them we propose a new diffusion-based Fourie…

Cited by 5SourcePDFScholar
2023

Document Image Shadow Removal Guided by Color-Aware Background

CVPR 2023poster

Existing works on document image shadow removal mostly depend on learning and leveraging a constant background (the color of the paper) from the image. However, the constant background is less representative and frequently ignores other background colors, such as the printed colors, resulting in dis…

2023

Learning Long-Range Information with Dual-Scale Transformers for Indoor Scene Completion

ICCV 2023poster

Due to the limited resolution of 3D sensors and the inevitable mutual occlusion between objects, 3D scans of real scenes are commonly incomplete. Previous scene completion methods struggle to capture long-range spatial feature, resulting in unsatisfactory completion results. To alleviate the pro…

Cited by 3PDFScholar
2023

NeTO:Neural Reconstruction of Transparent Objects with Self-Occlusion Aware Refraction-Tracing

ICCV 2023poster

We present a novel method called NeTO, for capturing the 3D geometry of solid transparent objects from 2D images via volume rendering. Reconstructing transparent objects is a very challenging task, which is ill-suited for general-purpose reconstruction techniques due to the specular light transport…

Cited by 12PDFScholar
2023

Towards High-Quality Specular Highlight Removal by Leveraging Large-Scale Synthetic Data

ICCV 2023poster

This paper aims to remove specular highlights from a single object-level image. Although previous methods have made some progresses, their performance remains somewhat limited, particularly for real images with complex specular highlights. To this end, we propose a three-stage network to address the…

Cited by 14PDFcodeScholar
2022

DGECN: A Depth-Guided Edge Convolutional Network for End-to-End 6D Pose Estimation

CVPR 2022poster

Monocular 6D pose estimation is a fundamental task in computer vision. Existing works often adopt a twostage pipeline by establishing correspondences and utilizing a RANSAC algorithm to calculate 6 degrees-of-freedom (6DoF) pose. Recent works try to integrate differentiable RANSAC algorithms to achi…

Cited by 39PDFScholar
2022

Deep Image-Based Illumination Harmonization

CVPR 2022poster

Integrating a foreground object into a background scenewith illumination harmonization is an important but chal-lenging task in computer vision and augmented reality community. Existing methods mainly focus on foreground andbackground appearance consistency or the foreground object shadow generation…

Cited by 18PDFcodeScholar
2022

Video Shadow Detection via Spatio-Temporal Interpolation Consistency Training

CVPR 2022poster

It is challenging to annotate large-scale datasets for supervised video shadow detection methods. Using a model trained on labeled images to the video frames directly may lead to high generalization error and temporal inconsistent results. In this paper, we address these challenges by proposing a Sp…

Cited by 21PDFcodeScholar
2021

A Multi-Task Network for Joint Specular Highlight Detection and Removal

CVPR 2021poster

Specular highlight detection and removal are fundamental and challenging tasks. Although recent methods achieve promising results on the two tasks by supervised training on synthetic training data, they are typically solely designed for highlight detection or removal, and their performance usually d…

Cited by 100PDFcodeScholar
2020

ARShadowGAN: Shadow Generative Adversarial Network for Augmented Reality in Single Light Scenes

CVPR 2020poster

Generating virtual object shadows consistent with the real-world environment shading effects is important but challenging in computer vision and augmented reality applications. To address this problem, we propose an end-to-end Generative Adversarial Network for shadow generation named ARShadowGAN fo…

Cited by 106PDFcodeScholar
2020

Detail Preserved Point Cloud Completion via Separated Feature Aggregation

ECCV 2020poster

Point cloud shape completion is a challenging problem in 3D vision and robotics. Existing learning-based frameworks leverage encoder-decoder architectures to recover the complete shape from a compactly encoded global feature vector. Though the global feature can approximately represent the overall l…

2019

ARGAN: Attentive Recurrent Generative Adversarial Network for Shadow Detection and Removal

ICCV 2019poster

In this paper we propose an attentive recurrent generative adversarial network (ARGAN) to detect and remove shadows in an image. The generator consists of multiple progressive steps. At each step a shadow attention detector is firstly exploited to generate an attention map which specifies shadow reg…

Cited by 174PDFScholar
2019

PCAN: 3D Attention Map Learning Using Contextual Information for Point Cloud Based Retrieval

CVPR 2019poster

Point cloud based retrieval for place recognition is an emerging problem in vision field. The main challenge is how to find an efficient way to encode the local features into a discriminative global descriptor. In this paper, we propose a Point Contextual Attention Network (PCAN), which can predict…

Cited by 266PDFcodeScholar
2018

Texture Mapping for 3D Reconstruction With RGB-D Sensor

CVPR 2018poster

Acquiring realistic texture details for 3D models is important in 3D reconstruction. However, the existence of geometric errors, caused by noisy RGB-D sensor data, always makes the color images cannot be accurately aligned onto reconstructed 3D models. In this paper, we propose a global-to-local cor…

Cited by 102SourcePDFScholar
2017

Distinguishing the Indistinguishable: Exploring Structural Ambiguities via Geodesic Context

CVPR 2017spotlight

A perennial problem in structure from motion (SfM) is visual ambiguity posed by repetitive structures. Recent disambiguating algorithms infer ambiguities mainly via explicit background context, thus face limitations in highly ambiguous scenes which are visually indistinguishable. Instead of analyzin…

Cited by 38PDFcodeScholar