← Search

Xiao Bai

27 accepted papers

2026

CoLoR: The Devil is in Scene Coordinate Regression for Large-Scale Visual Localization

CVPR 2026

Scene Coordinate Regression (SCR) has emerged as a memory-efficient paradigm for visual localization. While SCR has demonstrated performance comparable to classic feature matching based approaches in small-scale scenes, it has consistently underperformed in large-scale environments. Large-scale loca

Cited by 0SourceScholar
2026

FoundationSLAM: Unleashing the Power of Depth Foundation Models for End-to-End Dense Visual SLAM

AAAI 2026technical

We present FoundationSLAM, a learning-based monocular dense SLAM system that addresses the absence of geometric consistency in previous flow-based approaches for accurate and robust tracking and mapping. Our core idea is to bridge flow estimation with geometric reasoning by leveraging the guidance f

Cited by 0SourcePDFScholar
2026

MTAttack: Multi-Target Backdoor Attacks Against Large Vision-Language Models

AAAI 2026technical

Recent advances in Large Visual Language Models (LVLMs) have demonstrated impressive performance across various vision-language tasks by leveraging large-scale image-text pretraining and instruction tuning. However, the security vulnerabilities of LVLMs have become increasingly concerning, particula

Cited by 0SourcePDFScholar
2026

Revisiting Photometric Ambiguity for Accurate Gaussian-Splatting Surface Reconstruction

ICML 2026poster

Surface reconstruction with differentiable rendering has achieved impressive performance in recent years, yet the pervasive photometric ambiguities have strictly bottlenecked existing approaches. This paper presents AmbiSuR, a framework that explores an intrinsic solution upon Gaussian Splatting for…

Cited by 0SourceScholar
2026

SCE-SLAM: Scale-Consistent Monocular SLAM via Scene Coordinate Embeddings

CVPR 2026

Monocular visual SLAM enables 3D reconstruction from internet video and autonomous navigation on resource-constrained platforms, yet suffers from scale drift, i.e., the gradual divergence of estimated scale over long sequences. Existing frame-to-frame methods achieve real-time performance through lo

Cited by 0SourceScholar
2026

SparseSurf: Sparse-View 3D Gaussian Splatting for Surface Reconstruction

AAAI 2026technical

Recent advances in optimizing Gaussian Splatting for scene geometry have enabled efficient reconstruction of detailed surfaces from images. However, when input views are sparse, such optimization is prone to overfitting, leading to suboptimal reconstruction quality. Existing approaches address this

Cited by 0SourcePDFScholar
2025

Auxiliary Prompt Tuning of Vision-Language Models for Few-Shot Out-of-Distribution Detection

ICCV 2025poster

Recent advancements in CLIP-based out-of-distribution (OOD) detection have shown promising results via regularization on prompt tuning, leveraging background features extracted from a few in-distribution (ID) samples as proxies for OOD features.However, these methods suffer from an inherent limitati…

2025

Eve3D: Elevating Vision Models for Enhanced 3D Surface Reconstruction via Gaussian Splatting

NeurIPS 2025poster

We present Eve3D, a novel framework for dense 3D reconstruction based on 3D Gaussian Splatting (3DGS). While most existing methods rely on imperfect priors derived from pre-trained vision models, Eve3D fully leverages these priors by jointly optimizing both them and the 3DGS backbone. This joint opt…

Cited by 0SourceScholar
2025

GeoSVR: Taming Sparse Voxels for Geometrically Accurate Surface Reconstruction

NeurIPS 2025spotlight

Reconstructing accurate surfaces with radiance fields has achieved remarkable progress in recent years. However, prevailing approaches, primarily based on Gaussian Splatting, are increasingly constrained by representational bottlenecks. In this paper, we introduce GeoSVR, an explicit voxel-based fra…

Cited by 0SourcecodeScholar
2025

InsTaG: Learning Personalized 3D Talking Head from Few-Second Video

CVPR 2025poster

Despite exhibiting impressive performance in synthesizing lifelike personalized 3D talking heads, prevailing methods based on radiance fields suffer from high demands for training data and time for each new identity. This paper introduces InsTaG, a 3D talking head synthesis framework that allows a f…

2025

Revisiting Continual Ultra-fine-grained Visual Recognition with Pre-trained Models

IJCAI 2025

Continual ultra-fine-grained visual recognition (C-UFG) aims to continuously learn to categorize the increasing number of cultivates (VC-UFG) and consistently recognize crops across reproductive stages (HC-UFG), which is a fundamental goal of intelligent agriculture. Despite the progress made in gen

2025

View-aware Decomposition and Unification for Fast Ground-to-Aerial Person Search

IROS 2025

Ground-to-aerial person search leverages cooperative efforts between unmanned aerial vehicles (UAV) and ground surveillance cameras to locate person individuals. Despite the progress made by recent works, the impact of the discrepancy between the two views is underestimated. This limits the overall

Cited by 0SourcecodeScholar
2024

DNGaussian: Optimizing Sparse-View 3D Gaussian Radiance Fields with Global-Local Depth Normalization

CVPR 2024poster

Radiance fields have demonstrated impressive performance in synthesizing novel views from sparse input views yet prevailing methods suffer from high training costs and slow inference speed. This paper introduces DNGaussian a depth-regularized framework based on 3D Gaussian radiance fields offering r…

2024

Learning Transferable Negative Prompts for Out-of-Distribution Detection

CVPR 2024poster

Existing prompt learning methods have shown certain capabilities in Out-of-Distribution (OOD) detection but the lack of OOD images in the target dataset in their training can lead to mismatches between OOD images and In-Distribution (ID) categories resulting in a high false positive rate. To address…

2024

Long-Tailed Out-of-Distribution Detection via Normalized Outlier Distribution Adaptation

NeurIPS 2024poster

One key challenge in Out-of-Distribution (OOD) detection is the absence of ground-truth OOD samples during training. One principled approach to address this issue is to use samples from external datasets as outliers ($\textit{i.e.}$, pseudo OOD samples) to train OOD detectors. However, we find emp…

2024

Out-of-Distribution Detection in Long-Tailed Recognition with Calibrated Outlier Class Learning

AAAI 2024technical

Existing out-of-distribution (OOD) methods have shown great success on balanced datasets but become ineffective in long-tailed recognition (LTR) scenarios where 1) OOD samples are often wrongly classified into head classes and/or 2) tail-class samples are treated as OOD samples. To address these iss…

2024

Robust Synthetic-to-Real Transfer for Stereo Matching

CVPR 2024poster

With advancements in domain generalized stereo matching networks models pre-trained on synthetic data demonstrate strong robustness to unseen domains. However few studies have investigated the robustness after fine-tuning them in real-world scenarios during which the domain generalization ability ca…

2024

Simple Image-Level Classification Improves Open-Vocabulary Object Detection

AAAI 2024technical

Open-Vocabulary Object Detection (OVOD) aims to detect novel objects beyond a given set of base categories on which the detection model is trained. Recent OVOD methods focus on adapting the image-level pre-trained vision-language models (VLMs), such as CLIP, to a region-level object detection task v…

2023

Efficient Region-Aware Neural Radiance Fields for High-Fidelity Talking Portrait Synthesis

ICCV 2023poster

This paper presents ER-NeRF, a novel conditional Neural Radiance Fields (NeRF) based architecture for talking portrait synthesis that can concurrently achieve fast convergence, real-time rendering, and state-of-the-art performance with small model size. Our idea is to explicitly exploit the unequal…

Cited by 84PDFcodeScholar
2022

Revisiting Domain Generalized Stereo Matching Networks From a Feature Consistency Perspective

CVPR 2022poster

Despite recent stereo matching networks achieving impressive performance given sufficient training data, they suffer from domain shifts and generalize poorly to unseen domains. We argue that maintaining feature consistency between matching pixels is a vital factor for promoting the generalization ca…

Cited by 79PDFcodeScholar
2022

Where to Focus: Investigating Hierarchical Attention Relationship for Fine-Grained Visual Classification

ECCV 2022poster

"Object categories are often grouped into a multi-granularity taxonomic hierarchy. Classifying objects at coarser-grained hierarchy requires global and common characteristics, while finer-grained hierarchy classification relies on local and discriminative features. Therefore, humans should also subc…

2021

Goal-Oriented Gaze Estimation for Zero-Shot Learning

CVPR 2021poster

Zero-shot learning (ZSL) aims to recognize novel classes by transferring semantic knowledge from seen classes to unseen classes. Since semantic knowledge is built on attributes shared between different classes, which are highly local, strong prior for localization of object attribute is beneficial f…

Cited by 172PDFcodeScholar
2021

Occluded Person Re-Identification With Single-Scale Global Representations

ICCV 2021poster

Occluded person re-identification (ReID) aims at re-identifying occluded pedestrians from occluded or holistic images taken across multiple cameras. Current state-of-the-art (SOTA) occluded ReID models rely on some auxiliary modules, including pose estimation, feature pyramid and graph matching modu…

Cited by 64PDFScholar
2021

Select, Extract and Generate: Neural Keyphrase Generation with Layer-wise Coverage Attention

ACL 2021long

Natural language processing techniques have demonstrated promising results in keyphrase generation. However, one of the major challenges in neural keyphrase generation is processing long documents using deep neural networks. Generally, documents are truncated before given as inputs to neural network…

2020

Self-Trained Deep Ordinal Regression for End-to-End Video Anomaly Detection

CVPR 2020poster

Video anomaly detection is of critical practical importance to a variety of real applications because it allows human attention to be focused on events that are likely to be of interest, in spite of an otherwise overwhelming volume of video. We show that applying self-trained deep ordinal regression…

Cited by 318PDFScholar