← Search

Yihong Wu

31 accepted papers

2026

AING-SLAM: Accurate Implicit Neural Geometry-Aware SLAM With Appearance and Semantics via History-Guided Optimization

RA-L 2026

In range-based SLAM systems, localization accuracy depends on the quality of geometric maps. Sparse LiDAR scans and noisy depth from RGB-D sensors often yield incomplete or inaccurate reconstructions that degrade pose estimation. Appearance and semantic cues, readily available from onboard RGB and p

Cited by 0SourceScholar
2026

MORE-STEM: Long-Short MemOry REcall and Spatio-TEmporal Consistency Model for Query-Driven 3D/4D Point Cloud Segmentation

CVPR 2026

Current query-driven 3D understanding methods are constrained to static point clouds, limiting their ability to reason about dynamic scenes. To bridge this gap, we propose MORE-STEM, a unified framework for Long-Short MemOry REcall and Spatio-TEmporal Consistency Model in Query-Driven 3D/4D Point Cl

Cited by 0SourceScholar
2026

MR-COSMO: Visual-Text Memory Recall and Direct CrOSs-MOdal Alignment Method for Query-Driven 3D Segmentation

AAAI 2026technical

The rapid advancement of vision-language models (VLMs) in 3D domains has accelerated research in text-query-guided point cloud processing, though existing methods underperform in point-level segmentation due to inadequate 3D-text alignment that limits local feature-text context linking. To address t

Cited by 0SourcePDFScholar
2026

Sparse3DPR: Training-Free 3D Hierarchical Scene Parsing and Task-Adaptive Subgraph Reasoning from Sparse RGB Views

AAAI 2026technical

Recently, large language models (LLMs) have been explored widely for 3D scene understanding. Among them, training-free approaches are gaining attention for their flexibility and generalization over training-based methods. However, they typically struggle with accuracy and efficiency in practical dep

Cited by 0SourcePDFScholar
2025

3D-SLNR: A Super Lightweight Neural Representation for Large-scale 3D Mapping

CVPR 2025poster

We propose 3D-SLNR, a new and ultra-lightweight neural representation with outstanding performance for large-scale 3D mapping. The representation defines a global signed distance function (SDF) in near-surface space based on a set of band-limited local SDFs anchored at support points sampled from po…

Cited by 0SourcePDFScholar
2025

DOGE: An Extrinsic Orientation and Gyroscope Bias Estimation for Visual-Inertial Odometry Initialization

ICRA 2025

Most existing visual-inertial odometry (VIO) initialization methods rely on accurate pre-calibrated extrinsic parameters. However, during long-term use, irreversible structural deformation caused by temperature changes, mechanical squeezing, etc. will cause changes in extrinsic parameters, especiall

Cited by 3SourceScholar
2025

FEAST-Mamba: FEAture and SpaTial Aware Mamba Network with Bidirectional Orthogonal Fusion for Cross-Modal Point Cloud Segmentation

AAAI 2025technical

Point cloud segmentation has a wide range of applications in autonomous driving, augmented reality and virtual reality. Multi-modal fusion strategies have received increasing attention in point cloud segmentation recently. Despite the success, existing methods usually generate unnecessary informatio…

Cited by 0SourcePDFScholar
2025

Floorplan-SLAM: A Real-Time, High-Accuracy, and Long-Term Multi-Session Point-Plane SLAM for Efficient Floorplan Reconstruction

IROS 2025

Floorplan reconstruction provides structural priors essential for reliable indoor robot navigation and high-level scene understanding. However, existing approaches either require time-consuming offline processing with a complete map, or rely on expensive sensors and substantial computational resourc

Cited by 1SourceScholar
2025

Hi-Gaussian: Hierarchical Gaussians under Normalized Spherical Projection for Single-View 3D Reconstruction

ICCV 2025poster

Single-view 3D reconstruction is a fundamental problem in computer vision, having a significant impact on downstream tasks such as autonomous driving, virtual reality and augmented reality. However, existing single-view reconstruction methods are unable to reconstruct the regions outside the input f…

Cited by 0SourcePDFScholar
2025

Maximum Clique-Based Floorplan Association for Robust Multi-Session Stereo SLAM in Challenging Indoor Environments

IROS 2025

Existing multi-session visual simultaneous localization and mapping (SLAM) systems struggle severely to achieve robust localization and map merging under extreme viewpoint and illumination variations, particularly when handling completely opposite viewpoints and drastic day-night lighting changes. T

Cited by 0SourceScholar
2025

PGD-VIO: A Plane-Aided RGB-D Inertial Odometry With Graph-Based Drift Suppression

RA-L 2025

Generally, high-level features provide more geometrical information compared to point features, which can be exploited to further constrain motions. Planes are commonplace in man-made environments, offering an active means to reduce drift, due to their extensive spatial and temporal observability. T

Cited by 2SourceScholar
2024

Exploring the Best Practices of Query Expansion with Large Language Models

EMNLP 2024finding

Large Language Models (LLMs) are foundational in language technologies, particularly in information retrieval (IR). In this paper, we thoroughly explore the best practice of leveraging LLMs for query expansion. To this end, we introduce a training-free, straightforward yet effective framework called…

2024

RSS: Robust Stereo SLAM With Novel Extraction and Full Exploitation of Plane Features

RA-L 2024

Planar structures, prevalent in man-made environments, can be observed by a camera for significant periods of time due to their large spatial presence. These structures provide strong planar regularities for Simultaneous Localization and Mapping (SLAM) systems, facilitating long-term navigation. The

Cited by 12SourceScholar
2024

ViSTec: Video Modeling for Sports Technique Recognition and Tactical Analysis

AAAI 2024technical

The immense popularity of racket sports has fueled substantial demand in tactical analysis with broadcast videos. However, existing manual methods require laborious annotation, and recent attempts leveraging video perception models are limited to low-level annotations like ball trajectories, overloo…

2023

Accurate Implicit Neural Mapping With More Compact Representation in Large-Scale Scenes Using Ranging Data

RA-L 2023

Large-scale 3D mapping nowadays is a research hotspot in robotics. A greatly concerning issue is reconstructing high-accuracy maps in a hardware environment with limited memory. To address this problem, we propose a novel implicit neural mapping approach with higher accuracy and less memory. It firs

Cited by 13SourceScholar
2023

An Accurate Outlier Rejection Network With Higher Generalization Ability for Point Cloud Registration

RA-L 2023

Feature-based point cloud registration algorithms have gained more attention recently for their high robustness. Outlier rejection is a key step of such algorithms. With the development of deep learning, some of the learning-based outlier rejection methods have been proposed and implemented in vario

Cited by 9SourceScholar
2023

ConvGQR: Generative Query Reformulation for Conversational Search

ACL 2023long

In conversational search, the user’s real search intent for the current conversation turn is dependent on the previous conversation history. It is challenging to determine a good search query from the whole conversation context. To avoid the expensive re-training of the query encoder, most existing…

2023

Depth Estimation for a Single Omnidirectional Image with Reversed-Gradient Warming-up Thresholds Discriminator

ICASSP 2023accepted

Depth estimation for single image using deep learning requires a large labelled depth dataset with various scenes for training. However, currently published omnidirectional depth datasets cover limited types of scenes and are not suitable for depth estimation for various real-world scenes. With the…

Cited by 0SourceScholar
2023

MoqaGPT : Zero-Shot Multi-modal Open-domain Question Answering with Large Language Model

EMNLP 2023long findings

Multi-modal open-domain question answering typically requires evidence retrieval from databases across diverse modalities, such as images, tables, passages, etc. Even Large Language Models (LLMs) like GPT-4 fall short in this task. To enable LLMs to tackle the task in a zero-shot manner, we introduc…

Cited by 0SourcecodeScholar
2023

PLPL-VIO: A Novel Probabilistic Line Measurement Model for Point-Line-Based Visual-Inertial Odometry

IROS 2023poster

Point and line features are complementary in Visual-Inertial Odometry (VIO) or Visual-Inertial Simultaneous Localization And Mapping (VI-SLAM) systems. The advantage of combining these two types of features relies on their proper weighting in the cost function, usually set by their uncertainty. Comp…

Cited by 5SourceScholar
2022

Structural Regularity Aided Visual-Inertial Odometry With Novel Coordinate Alignment and Line Triangulation

RA-L 2022

Man-made buildings exhibit structural regularity, which can provide strongly geometrical constraints for Visual-Inertial Odometry (VIO) systems. To make full use of the structural information, we propose a new structural regularity aided VIO with novel coordinate alignment and line triangulation und

Cited by 17SourceScholar
2021

A Flexible and Efficient Loop Closure Detection Based on Motion Knowledge

ICRA 2021poster

Loop closure detection (LCD) is an essential module for simultaneous localization and mapping (SLAM), which can correct accumulated errors after long-term explorations. The widely used bag-of-words (BoW) model can not satisfy well the requirements of both low time consumption and high accuracy for a…

Cited by 5SourceScholar
2021

A Point-Line VIO System With Novel Feature Hybrids and With Novel Line Predicting-Matching

RA-L 2021

Weak texture and motion blur are always challenging problems for visual-inertial odometry (VIO) systems. To improve accuracy of VIO systems in the challenging scenes, we propose a point-line-based VIO system with novel feature hybrids and with novel predicting-matching for long line track. Point-lin

Cited by 28SourceScholar
2021

Highly Efficient Line Segment Tracking with an IMU-KLT Prediction and a Convex Geometric Distance Minimization

ICRA 2021poster

Line segment features become popular in SLAM community. Usually, line-based SLAM systems utilize local appearance descriptors for line segment tracking. However, traditional descriptor-based line segment tracking algorithms suffer from the problem that accuracy and speed cannot be possessed simultan…

Cited by 20SourceScholar
2020

Spectral Graph Matching and Regularized Quadratic Relaxations: Algorithm and Theory

ICML 2020poster

Graph matching, also known as network alignment, aims at recovering the latent vertex correspondence between two unlabeled, edge-correlated weighted graphs. To tackle this task, we propose a spectral method, GRAph Matching by Pairwise eigen-Alignments (GRAMPA), which first constructs a similarity ma…

Cited by 55SourcePDFScholar
2019

FMD Stereo SLAM: Fusing MVG and Direct Formulation Towards Accurate and Fast Stereo SLAM

ICRA 2019poster

We propose a novel stereo visual SLAM framework considering both accuracy and speed at the same time. The framework makes full use of the advantages of key-feature-based multiple view geometry (MVG) and direct-based formulation. At the front-end, our system performs direct formulation and constant m…

Cited by 18SourceScholar
2018

Data Amplification: A Unified and Competitive Approach to Property Estimation

NeurIPS 2018poster

Estimating properties of discrete distributions is a fundamental problem in statistical learning. We design the first unified, linear-time, competitive, property estimator that for a wide class of properties and for all underlying distributions uses just 2n samples to achieve the performance attaine…

Cited by 30SourcePDFScholar
2018

Entropy Rate Estimation for Markov Chains with Large State Space

NeurIPS 2018spotlight

Entropy estimation is one of the prototypical problems in distribution property testing. To consistently estimate the Shannon entropy of a distribution on $S$ elements with independent samples, the optimal sample complexity scales sublinearly with $S$ as $\Theta(\frac{S}{\log S})$ as shown by Valian…

Cited by 22SourcePDFScholar