← Search

Xuanhong Chen

19 accepted papers

2026

FastAvatar: Towards Unified and Fast 3D Avatar Reconstruction with Large Gaussian Reconstruction Transformers

ICLR 2026poster

Despite significant progress in 3D avatar reconstruction, it still faces challenges such as high time complexity, sensitivity to data quality, and low data utilization. We propose~\textbf{FastAvatar}, a feedforward 3D avatar framework capable of flexibly leveraging diverse daily recordings (e.g., a…

Cited by 0SourcecodeScholar
2026

On the Computational Limits of AI4S-RL : A Unified $\varepsilon$-$N$ Analysis

ICLR 2026poster

Recent work increasingly adopts AI for Science (AI4S) models to replace expensive PDE solvers as simulation environments for reinforcement learning (RL), enabling faster training in complex physical control tasks. However, using approximate simulators introduces modeling errors that affect the learn…

Cited by 0SourceScholar
2025

RAGDiffusion: Faithful Cloth Generation via External Knowledge Assimilation

ICCV 2025poster

Standard clothing asset generation involves restoring forward-facing flat-lay garment images displayed on a clear background by extracting clothing information from diverse real-world contexts, which presents significant challenges due to highly standardized structure sampling distributions and clot…

Cited by 0SourcePDFScholar
2025

ShoeFit: A New Dataset and Dual-image-stream DiT Framework for Virtual Footwear Try-On

NeurIPS 2025poster

Virtual footwear try-on (VFTON), a critical yet underexplored area in virtual try-on (VTON), aims to synthesize faithful try-on results given diverse footwear and model images while maintaining 3D consistency and texture authenticity. Unlike conventional garment-focused VTON methods, VFTON present…

Cited by 0SourceScholar
2025

SinGS: Animatable Single-Image Human Gaussian Splats with Kinematic Priors

CVPR 2025poster

Despite significant advances in accurately estimating geometry in contemporary single-image 3D human reconstruction, creating a high-quality, efficient, and animatable 3D avatar remains an open challenge. Two key obstacles persist: incomplete observation and inconsistent 3D priors. To address these…

2024

AnyFit: Controllable Virtual Try-on for Any Combination of Attire Across Any Scenario

NeurIPS 2024poster

While image-based virtual try-on has made significant strides, emerging approaches still fall short of delivering high-fidelity and robust fitting images across various scenarios, as their models suffer from issues of ill-fitted garment styles and quality degrading during the training process, not t…

Cited by 7SourcePDFScholar
2024

FocalDreamer: Text-Driven 3D Editing via Focal-Fusion Assembly

AAAI 2024technical

While text-3D editing has made significant strides in leveraging score distillation sampling, emerging approaches still fall short in delivering separable, precise and consistent outcomes that are vital to content creation. In response, we introduce FocalDreamer, a framework that merges base shape w…

Cited by 56SourcePDFScholar
2024

Intrinsic Phase-Preserving Networks for Depth Super Resolution

AAAI 2024technical

Depth map super-resolution (DSR) plays an indispensable role in 3D vision. We discover an non-trivial spectral phenomenon: the components of high-resolution (HR) and low-resolution (LR) depth maps manifest the same intrinsic phase, and the spectral phase of RGB is a superset of them, which suggests…

2024

Toward Tiny and High-quality Facial Makeup with Data Amplify Learning

ECCV 2024poster

"Contemporary makeup approaches primarily hinge on unpaired learning paradigms, yet they grapple with the challenges of inaccurate supervision (e.g., face misalignment) and sophisticated facial prompts (including face parsing, and landmark detection). These challenges prohibit low-cost deployment of…

2024

Towards High-fidelity Artistic Image Vectorization via Texture-Encapsulated Shape Parameterization

CVPR 2024poster

We develop a novel vectorized image representation scheme accommodating both shape/geometry and texture in a decoupled way particularly tailored for reconstruction and editing tasks of artistic/design images such as Emojis and Cliparts. In the heart of this representation is a set of sparsely and un…

Cited by 1SourcePDFScholar
2023

Deep Arbitrary-Scale Image Super-Resolution via Scale-Equivariance Pursuit

CVPR 2023poster

The ability of scale-equivariance processing blocks plays a central role in arbitrary-scale image super-resolution tasks. Inspired by this crucial observation, this work proposes two novel scale-equivariant modules within a transformer-style framework to enhance arbitrary-scale image super-resolutio…

2023

Generalized Deep 3D Shape Prior via Part-Discretized Diffusion Process

CVPR 2023poster

We develop a generalized 3D shape generation prior model, tailored for multiple 3D tasks including unconditional shape generation, point cloud completion, and cross-modality shape generation, etc. On one hand, to precisely capture local fine detailed shape information, a vector quantized variational…

2023

Learning Continuous Depth Representation via Geometric Spatial Aggregator

AAAI 2023technical

Depth map super-resolution (DSR) has been a fundamental task for 3D computer vision. While arbitrary scale DSR is a more realistic setting in this scenario, previous approaches predominantly suffer from the issue of inefficient real-numbered scale upsampling. To explicitly address this issue, we pro…

2023

Omni Aggregation Networks for Lightweight Image Super-Resolution

CVPR 2023poster

While lightweight ViT framework has made tremendous progress in image super-resolution, its uni-dimensional self-attention modeling, as well as homogeneous aggregation scheme, limit its effective receptive field (ERF) to include more comprehensive interactions from both spatial and channel dimension…

2022

Bi-volution: A Static and Dynamic Coupled Filter

AAAI 2022technical

Dynamic convolution has achieved significant gain in performance and computational complexity, thanks to its powerful representation capability given limited filter number/layers. However, SOTA dynamic convolution operators are sensitive to input noises (e.g., Gaussian noise, shot noise, e.t.c.) an…

2022

RainNet: A Large-Scale Imagery Dataset and Benchmark for Spatial Precipitation Downscaling

NeurIPS 2022accept

AI-for-science approaches have been applied to solve scientific problems (e.g., nuclear fusion, ecology, genomics, meteorology) and have achieved highly promising results. Spatial precipitation downscaling is one of the most important meteorological problem and urgently requires the participation of…

2021

Sketch Generation with Drawing Process Guided by Vector Flow and Grayscale

AAAI 2021technical

We propose a novel image-to-pencil translation method that could not only generate high-quality pencil sketches but also offer the drawing process. Existing pencil sketch algorithms are based on texture rendering rather than the direct imitation of strokes, making them unable to show the drawing pro…

2020

CooGAN: A Memory-Efficient Framework for High-Resolution Facial Attribute Editing

ECCV 2020poster

In contrast to great success of memory-consuming face editing methods at a low resolution, to manipulate high-resolution (HR) facial images, \ie, typically larger than $768^2$ pixels, with very limited memory is still challenging. This is due to the reasons of 1) intractable huge demand of memory; 2…