← Search

Zhihang Li

8 accepted papers

2026

Confidence-Guided Multi-Scale Aggregation for Sparse-View High-Resolution 3D Gaussian Splatting

CVPR 2026

Sparse-view 3D Gaussian Splatting (3DGS) reconstructs scenes using 3D Gaussians from sparse input views. Yet, this method is prone to overfitting, which is exacerbated at higher resolutions as the expanded dimensionality amplifies floating artifacts and reconstruction ambiguities. In this paper, we

Cited by 0SourceScholar
2026

One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer

CVPR 2026

Recent advances in diffusion models have greatly improved pose-driven character animation. However, existing methods are limited to spatially aligned reference-pose pairs with matched skeletal structures. Handling reference-pose misalignment remains unsolved. To address this, we present One-to-All A

Cited by 0SourcecodeScholar
2025

Diffusion-based Synthetic Data Generation for Visible-Infrared Person Re-Identification

AAAI 2025technical

The performance of models is intricately linked to the abundance of training data. In Visible-Infrared person Re-IDentification (VI-ReID) tasks, collecting and annotating large-scale images of each individual under various cameras and modalities is tedious, time-expensive, costly and must comply wit…

Cited by 0SourcePDFScholar
2025

Self-supervised ControlNet with Spatio-Temporal Mamba for Real-world Video Super-resolution

CVPR 2025poster

Existing diffusion-based video super-resolution (VSR) methods are susceptible to introducing complex degradations and noticeable artifacts into high-resolution videos due to their inherent randomness. In this paper, we propose a noise-robust real-world VSR framework by incorporating self-supervised…

Cited by 0SourcePDFScholar
2021

Neural Feature Search for RGB-Infrared Person Re-Identification

CVPR 2021poster

RGB-Infrared person re-identification (RGB-IR ReID) is a challenging cross-modality retrieval problem, which aims at matching the person-of-interest over visible and infrared camera views. Most existing works achieve performance gains through manually-designed feature selection modules, which often…

Cited by 251PDFcodeScholar
2020

GP-NAS: Gaussian Process Based Neural Architecture Search

CVPR 2020poster

Neural architecture search (NAS) advances beyond the state-of-the-art in various computer vision tasks by automating the designs of deep neural networks. In this paper, we aim to address three important questions in NAS: (1) How to measure the correlation between architectures and their performances…

Cited by 67PDFScholar
2019

ACFNet: Attentional Class Feature Network for Semantic Segmentation

ICCV 2019poster

Recent works have made great progress in semantic segmentation by exploiting richer context, most of which are designed from a spatial perspective. In contrast to previous works, we present the concept of class center which extracts the global context from a categorical perspective. This class-level…

Cited by 357PDFScholar
2018

IntroVAE: Introspective Variational Autoencoders for Photographic Image Synthesis

NeurIPS 2018poster

We present a novel introspective variational autoencoder (IntroVAE) model for synthesizing high-resolution photographic images. IntroVAE is capable of self-evaluating the quality of its generated samples and improving itself accordingly. Its inference and generator models are jointly trained in an i…

Cited by 356SourcePDFScholar