← Search

Shibiao Xu

22 accepted papers

2026

DialogueVPR: Towards Conversational Visual Place Recognition

CVPR 2026

Inspired by how humans communicate spatial information, language-guided geo-localization has gained significant traction for its intuitive and practical value. Despite this progress, most methods still rely on a static, one-shot retrieval paradigm, which fails to handle the ambiguity and incompleten

Cited by 0SourcecodeScholar
2026

RealNet: Efficient and Unsupervised Detection of AI-Generated Images via Real-Only Representation Learning

AAAI 2026technical

Detecting AI-generated images remains a persistent challenge, as existing detectors often struggle to generalize to forgeries produced by previously unseen generative models. This generalization gap mainly stems from entanglement with semantic content and overfitting to model-specific artifacts. Mor

Cited by 0SourcePDFScholar
2026

SAGE: Spatial-visual Adaptive Graph Exploration for Efficient Visual Place Recognition

ICLR 2026poster

Visual Place Recognition (VPR) requires robust retrieval of geotagged images despite large appearance, viewpoint, and environmental variation. Prior methods focus on descriptor fine-tuning or fixed sampling strategies yet neglect the dynamic interplay between spatial context and visual similarity d…

Cited by 0SourcecodeScholar
2025

3D-MoRe: Unified Modal-Contextual Reasoning for Embodied Question Answering

IROS 2025

With the growing need for diverse and scalable data in indoor scene tasks, such as question answering and dense captioning, we propose 3D-MoRe, a novel paradigm designed to generate large-scale 3D-language datasets by lever-aging the strengths of foundational models. The framework integrates key com

Cited by 12SourcecodeScholar
2025

Complementary Information Guided Occupancy Prediction via Multi-Level Representation Fusion

ICRA 2025

Camera-based occupancy prediction is a main-stream approach for 3D perception in autonomous driving, aiming to infer complete 3D scene geometry and semantics from 2D images. Almost existing methods focus on improving performance through structural modifications, such as lightweight backbones and com

Cited by 1SourceScholar
2025

DiffusionIMU: Diffusion-Based Inertial Navigation with Iterative Motion Refinement

IJCAI 2025

Inertial navigation enables self-contained localization using only Inertial Measurement Units (IMUs), making it widely applicable in various domains such as navigation, augmented reality, and robotics. However, existing methods suffer from drift accumulation due to the sensor noise and difficulty ca

Cited by 0SourcePDFScholar
2025

Focus on Local: Finding Reliable Discriminative Regions for Visual Place Recognition

AAAI 2025technical

Visual Place Recognition (VPR) is aimed at predicting the location of a query image by referencing a database of geotagged images. For VPR task, often fewer discriminative local regions in an image produce important effects while mundane background regions do not contribute or even cause perceptual…

2025

Novel View Synthesis Under Large-Deviation Viewpoint for Autonomous Driving

AAAI 2025technical

Novel view synthesis is a critical task in autonomous driving. Although 3D Gaussian Splatting (3D-GS) has shown success in generating novel views, it faces challenges in maintaining high-quality rendering when viewpoints deviate significantly from the training set. This difficulty primarily stems fr…

Cited by 0SourcePDFScholar
2025

SCS: Spatially Consistent Self-Supervised approach for One-Shot Anatomical Landmark Detection

ICASSP 2025accepted

Landmark detection is essential in medical image analysis, serving as the foundation for many downstream tasks. In recent years, supervised anatomical landmark detection models have achieved remarkable success, but typically require large amounts of labeled data for training, which is challenging to…

Cited by 0SourceScholar
2024

DefFusion: Deformable Multimodal Representation Fusion for 3D Semantic Segmentation

ICRA 2024poster

The complementarity between camera and LiDAR data makes fusion methods a promising approach to improve 3D semantic segmentation performance. Recent transformer-based methods have also demonstrated superiority in segmentation. However, multimodal solutions incorporating transformers are underexplored…

Cited by 8SourceScholar
2024

QAGait: Revisit Gait Recognition from a Quality Perspective

AAAI 2024technical

Gait recognition is a promising biometric method that aims to identify pedestrians from their unique walking patterns. Silhouette modality, renowned for its easy acquisition, simple structure, sparse representation, and convenient modeling, has been widely employed in controlled in-the-lab research.…

2024

Spectral Prompt Tuning: Unveiling Unseen Classes for Zero-Shot Semantic Segmentation

AAAI 2024technical

Recently, CLIP has found practical utility in the domain of pixel-level zero-shot segmentation tasks. The present landscape features two-stage methodologies beset by issues such as intricate pipelines and elevated computational costs. While current one-stage approaches alleviate these concerns and…

2024

UnionFormer: Unified-Learning Transformer with Multi-View Representation for Image Manipulation Detection and Localization

CVPR 2024poster

We present UnionFormer a novel framework that integrates tampering clues across three views by unified learning for image manipulation detection and localization. Specifically we construct a BSFI-Net to extract tampering features from RGB and noise views achieving enhanced responsiveness to boundary…

Cited by 10SourcePDFScholar
2023

Self Correspondence Distillation for End-to-End Weakly-Supervised Semantic Segmentation

AAAI 2023technical

Efficiently training accurate deep models for weakly supervised semantic segmentation (WSSS) with image-level labels is challenging and important. Recently, end-to-end WSSS methods have become the focus of research due to their high training efficiency. However, current methods suffer from insuffici…

2023

Treating Pseudo-labels Generation as Image Matting for Weakly Supervised Semantic Segmentation

ICCV 2023poster

Generating accurate pseudo-labels under the supervision of image categories is a crucial step in Weakly Supervised Semantic Segmentation (WSSS). In this work, we propose a Mat-Label pipeline that provides a fresh way to treat WSSS pseudo-labels generation as an image matting task. By taking a trimap…

Cited by 29PDFcodeScholar
2022

DOMAINDESC: Learning Local Descriptors With Domain Adaptation

ICASSP 2022accepted

Robust and efficient local descriptor is crucial in a wide range of applications. In this paper, we propose a novel descriptor DomainDesc which is invariant as much as possible by learning local Descriptor with Domain adaptation. We design the feature-level domain adaptation loss to improve robustne…

Cited by 0SourceScholar
2022

GeoROS: Georeferenced Real-time Orthophoto Stitching with Unmanned Aerial Vehicle

IROS 2022poster

Simultaneous orthophoto stitching during the flight of Unmanned Aerial Vehicles (UAV) can greatly promote the practicability and instantaneity of diverse applications such as emergency disaster rescue, digital agriculture, and cadastral survey, which is of remarkable interest in aerial photogrammetr…

Cited by 3SourceScholar
2022

MTLDesc: Looking Wider to Describe Better

AAAI 2022technical

Limited by the locality of convolutional neural networks, most existing local features description methods only learn local descriptors with local information and lack awareness of global and surrounding spatial context. In this work, we focus on making local descriptors ``look wider to describe bet…

2021

A Periodic Frame Learning Approach for Accurate Landmark Localization in M-Mode Echocardiography

ICASSP 2021accepted

Anatomical landmark localization has been a key challenge for medical image analysis. Existing researches mostly adopt CNN as the main architecture for landmark localization while they are not applicable to process image modalities with periodic structure. In this paper, we propose a novel two-stage…

Cited by 0SourceScholar
2020

DenseFusion: Large-Scale Online Dense Pointcloud and DSM Mapping for UAVs

IROS 2020poster

With the rapidly developing unmanned aerial vehicles, the requirements of generating maps efficiently and quickly are increasing. To realize online mapping, we develop a real-time dense mapping framework named DenseFusion which can incrementally generates dense geo-referenced 3D point cloud, digital…

Cited by 11SourceScholar