← Search

Yufeng Yue

56 accepted papers

2026

Energy-GS: Image Energy-guided Pose Alignment Gaussian Splatting with redesigned pose gradient flow

CVPR 2026

High-quality 3D scene representation in radiance fields relies on accurate camera poses which are often difficult to acquire in real-world scenarios. An effective solution is to use RGB images for the joint optimization of radiance fields and camera poses, an approach that has been well explored in

Cited by 0SourcecodeScholar
2026

FilterGS: Traversal-Free Parallel Filtering and Adaptive Shrinking for Large-Scale LoD 3D Gaussian Splatting

CVPR 2026

3D Gaussian Splatting has revolutionized neural rendering with real-time performance. However, scaling this approach to large scenes using Level-of-Detail methods faces critical challenges: inefficient serial traversal consuming over 60% of rendering time, and redundant Gaussian-tile pairs that incu

Cited by 0SourcecodeScholar
2026

Learning a Unified Latent Action Space from Videos with Action-centric Cycle Consistency

CVPR 2026

Video data provides a rich source beyond expensive action-labeled data for advancing robot learning. Recent approaches have demonstrated promising potential in leveraging video data by learning latent actions for policy training. The latent action tokenizer encodes latent actions between successive

Cited by 0SourceScholar
2026

MCOO-SLAM: A Multi-Camera Omnidirectional Object SLAM System

RA-L 2026

Object-level SLAM offers structured and semantically meaningful environment representations, making it more interpretable and suitable for high-level robotic tasks. However, most existing approaches rely on RGB-D sensors or monocular views, which suffer from narrow fields of view, occlusion sensitiv

Cited by 2SourceScholar
2026

OmniMap: A General Mapping Framework Integrating Optics, Geometry, and Semantics

ICRA 2026poster

Robotic systems demand accurate and comprehensive 3D environment perception, requiring simultaneous capture of photo-realistic appearance (optical), precise layout shape (geometric), and open-vocabulary scene understanding (semantic). Existing methods typically achieve only partial fulfillment of th…

2026

OpenIN: Open-Vocabulary Instance-Oriented Navigation in Dynamic Domestic Environments

ICRA 2026poster

In daily domestic settings, frequently used objects like cups often have unfixed positions and multiple instances within the same category, and their carriers frequently change as well. As a result, it becomes challenging for a robot to efficiently navigate to a specific instance. To tackle this cha…

2026

PIPS: Planar Instance 3D Reconstruction Leveraging Planar Structural Priors

ICRA 2026poster

Planar structures, ubiquitous in man-made indoor environments, enable compact and accurate scene abstraction for various downstream tasks. Recent methods distill planar features into learning-based MVS geometries to obtain coherent 3D plane estimation from multi-view inputs. However, the lack of exp…

Cited by 0codeScholar
2026

Reliable LiDAR Loop Detection through Structural Descriptors and Semantic Graph Matching

ICRA 2026poster

Outdoor loop closure detection is essential for mitigating accumulated drift in SLAM and generating a global consistent map. Semantic graph matching methods utilize object-level topology for distinctive scene representation but rely on environments with rich and distinguishable objects. Moreover, ac…

Cited by 0codeScholar
2026

USER: A Unified and Extensible System for Online Real-World Policy Learning in Embodied AI

RSS 2026poster

Online policy learning directly in the physical world is a promising yet challenging direction for embodied intelligence. Unlike simulation, real-world systems cannot be arbitrarily accelerated, cheaply reset, or massively replicated, which makes scalable data collection, heterogeneous deployment, a…

Cited by 0SourceScholar
2026

Video2Robo: 3DGS-based Synthetic Data from One Video Enables Scalable Robot Learning

CVPR 2026

Scalable robot learning is hindered by the high cost of acquiring diverse, high-quality embodied data. Existing data generation approaches partially mitigate this issue but typically depend on hard-to-access hardware and labor-intensive manual effort, with limited generalization to diverse scene con

Cited by 0SourceScholar
2025

DroneSplat: 3D Gaussian Splatting for Robust 3D Reconstruction from In-the-Wild Drone Imagery

CVPR 2025highlight

Drones have become essential tools for reconstructing wild scenes due to their outstanding maneuverability. Recent advances in radiance field methods have achieved remarkable rendering quality, providing a new avenue for 3D reconstruction from drone imagery. However, dynamic distractors in wild env…

Cited by 2SourcePDFScholar
2025

GaussianGraph: 3D Gaussian-Based Scene Graph Generation for Open-World Scene Understanding

IROS 2025

Recent advancements in 3D Gaussian Splatting(3DGS) have significantly improved semantic scene understanding, enabling natural language queries to localize objects within a scene. However, existing methods primarily focus on embedding compressed CLIP features to 3D Gaussians, suffering from low objec

Cited by 6SourcecodeScholar
2025

GraphMimic: Graph-to-Graphs Generative Modeling from Videos for Policy Learning

CVPR 2025poster

Learning from demonstration is a powerful method for robotic skill acquisition. However, the significant expense of collecting such action-labeled robot data presents a major bottleneck. Video data, a rich data source encompassing diverse behavioral and physical knowledge, emerges as a promising alt…

Cited by 0SourcePDFScholar
2025

Hierarchical Autoregressive Modeling With Multi-Scale Refinement for Robot Policy Learning

RA-L 2025

While autoregressive models demonstrate remarkable success in text and image generation, their application to robot policies suffers from weak holistic comprehension, cumulative errors, and limited multi-modal modeling capabilities, particularly in long-horizon tasks or multi-modal scenarios. This p

Cited by 1SourceScholar
2025

High-Precision Object Pose Estimation Using Visual-Tactile Information for Dynamic Interactions in Robotic Grasping

ICRA 2025

In various robotic applications, understanding accurate object poses for robots is essential for high-precision tasks such as factory assembly or daily insertions. Tactile sensing, which compensates for visual information, offers rich texture-based or force-based data for object pose estimation. How

Cited by 0SourceScholar
2025

Human Demonstrations are Generalizable Knowledge for Robots

IROS 2025

Learning from human demonstrations is an emerging trend for designing intelligent robotic systems. However, previous methods typically regard videos as instructions, simply dividing videos into action sequences for robotic repetition, which pose obstacles to generalization to diverse tasks or object

Cited by 11SourceScholar
2025

LGSDF: Continual Global Learning of Signed Distance Fields Aided by Local Updating

RA-L 2025

Implicit reconstruction of ESDF (Euclidean Signed Distance Field) involves training a neural network to regress the signed distance from any point to the nearest obstacle, which has the advantages of lightweight storage and continuous querying. However, existing algorithms usually rely on conflictin

Cited by 5SourcecodeScholar
2025

ORA-NET: Enhancing Image Feature Matching through Oriented Overlapping Region Alignment

IROS 2025

Image feature matching is a fundamental task in computer vision. Existing local feature matching methods can establish robust correspondences between image pairs. However, these methods heavily rely on dense local image features, making them susceptible to significant perspective differences, charac

Cited by 0SourceScholar
2025

Open-RGBT: Open-Vocabulary RGB-T Zero-Shot Semantic Segmentation in Open-World Environments

ICRA 2025

Semantic segmentation is a critical technique for effective scene understanding. Traditional RGB-T semantic segmentation models often struggle to generalize across diverse scenarios due to their reliance on pretrained models and predefined categories. Recent advancements in Visual Language Models (V

Cited by 0SourcecodeScholar
2025

OpenGS-Fusion: Open-Vocabulary Dense Mapping with Hybrid 3D Gaussian Splatting for Refined Object-Level Understanding

IROS 2025

Recent advancements in 3D scene understanding have made significant strides in enabling interaction with scenes using open-vocabulary queries, particularly for VR/AR and robotic applications. Nevertheless, existing methods are hindered by rigid offline pipelines and the inability to provide precise

Cited by 4SourcecodeScholar
2025

OpenGS-SLAM: Open-Set Dense Semantic SLAM with 3D Gaussian Splatting for Object-Level Scene Understanding

ICRA 2025

Recent advancements in 3D Gaussian Splatting have significantly improved the efficiency and quality of dense semantic SLAM. However, previous methods are generally constrained by limited-category pre-trained classifiers and implicit semantic representation, which hinder their performance in open-set

Cited by 15SourcecodeScholar
2025

OpenIN: Open-Vocabulary Instance-Oriented Navigation in Dynamic Domestic Environments

RA-L 2025

In daily domestic settings, frequently used objects like <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">cups</i> often have unfixed positions and multiple instances within the same category, and their carriers also frequently change. As a result, it

Cited by 20SourcecodeScholar
2025

OpenMIGS: Multi-granularity Information-preserving Open-Vocabulary 3D Gaussian Splatting

IROS 2025

Open-vocabulary scene understanding is critical for robotics, yet existing 3D Gaussian Splatting (3DGS) methods rely on compressed feature embeddings, compromising semantic fidelity and fine-grained interpretation. Although utilizing uncompressed high-dimensional features offers a potential solution

Cited by 0SourcecodeScholar
2025

OpenMulti: Open-Vocabulary Instance-Level Multi-Agent Distributed Implicit Mapping

RA-L 2025

Multi-agent distributed collaborative mapping provides comprehensive and efficient representations for robots. However, existing approaches lack instance-level awareness and semantic understanding of environments, limiting their effectiveness for downstream applications. To address this issue, we pr

Cited by 3SourceScholar
2025

OpenObj: Open-Vocabulary Object-Level Neural Radiance Fields With Fine-Grained Understanding

RA-L 2025

In recent years, there has been a surge of interest in open-vocabulary 3D scene reconstruction facilitated by visual language models (VLMs), which showcase remarkable capabilities in open-set retrieval tasks. Although the semantic ambiguity of existing point-wise feature maps is alleviated by open-v

Cited by 12SourceScholar
2025

OpenObject-NAV: Open-Vocabulary Object-Oriented Navigation Based on Dynamic Carrier-Relationship Scene Graph

IROS 2025

In everyday life, frequently used objects like cups often have unfixed positions and multiple instances within the same category, and their carriers frequently change as well. As a result, it becomes challenging for a robot to efficiently navigate to a specific instance. To tackle this challenge, th

Cited by 3SourcecodeScholar
2025

OpenVox: Real-time Instance-level Open-vocabulary Probabilistic Voxel Representation

IROS 2025

In recent years, vision-language models (VLMs) have advanced open-vocabulary mapping, enabling mobile robots to simultaneously achieve environmental reconstruction and high-level semantic understanding. While integrated object cognition helps mitigate semantic ambiguity in point-wise feature maps, e

Cited by 4SourcecodeScholar
2025

PartGrasp: Generalizable Part-level Grasping via Semantic-Geometric Alignment

IROS 2025

The ability to perform generalizable and precise grasping on functional object parts is a prerequisite for robotic manipulation in open environments. Recent foundation models have demonstrated promising semantic correspondence capabilities in guiding robots to grasp similar parts across objects with

Cited by 0SourcecodeScholar
2025

SLOOP: Aligned Coordinate System-aided LiDAR LOOP Closure Detection based on Semantic Node Graph Matching

IROS 2025

Loop closure detection and pose estimation play a significant role in correcting odometry trajectories and generating globally consistent point cloud maps. Geometric feature descriptor methods neglect object-level spatial topology features, resulting in inadequate performance in loop closure detecti

Cited by 0SourcecodeScholar
2025

STEP Planner: Constructing cross-hierarchical subgoal tree as an embodied long-horizon task planner

IROS 2025

The ability to perform reliable long-horizon task planning is crucial for deploying robots in real-world environments. However, directly employing Large Language Models (LLMs) as action sequence generators often results in low success rates due to their limited reasoning ability for long-horizon emb

Cited by 4SourceScholar
2025

TASeg: Text-aware RGB-T Semantic Segmentation based on Fine-tuning Vision Foundation Models

IROS 2025

Reliable semantic segmentation of open environments is essential for intelligent systems, yet significant problems remain: 1) Existing RGB-T semantic segmentation models mainly rely on low-level visual features and lack high-level textual information, which struggle with accurate segmentation when c

Cited by 1SourceScholar
2024

Fast and Robust Point Cloud Registration with Tree-based Transformer

ICRA 2024poster

Point cloud registration is essential in computer vision and robotics. Recently, transformer-based methods have achieved advanced point cloud registration performance. However, the standard attention mechanism utilized in these methods considers many low-relevance points, and it has difficulty focus…

Cited by 2SourcecodeScholar
2024

LCP-Fusion: A Neural Implicit SLAM with Enhanced Local Constraints and Computable Prior

IROS 2024poster

Recently the dense Simultaneous Localization and Mapping (SLAM) based on neural implicit representation has shown impressive progress in hole filling and high-fidelity mapping. Nevertheless, existing methods either heavily rely on known scene bounds or suffer inconsistent reconstruction due to drift…

Cited by 0SourcecodeScholar
2024

MM4MM: Map Matching Framework for Multi-Session Mapping in Ambiguous and Perceptually-Degraded Environments

ICRA 2024poster

Multi-session mapping serves as the pre-requisite for autonomous robots to fulfill various long-term tasks (e.g., map updating, navigation, collaboration). However, it is challenging to implement multi-session mapping in enclosed or partially enclosed ambiguous environments (e.g., long corridors, in…

Cited by 0SourceScholar
2024

OpenGraph: Open-Vocabulary Hierarchical 3D Graph Representation in Large-Scale Outdoor Environments

RA-L 2024

Environment representations endowed with sophisticated semantics are pivotal for facilitating seamless interaction between robots and humans, enabling them to effectively carry out various tasks. Open-vocabulary representation, powered by Visual-Language models (VLMs), possesses inherent advantages,

Cited by 38SourcecodeScholar
2024

Robust Collaborative Perception against Temporal Information Disturbance

ICRA 2024poster

Collaborative perception facilitates a more comprehensive representation of the environment by leveraging complementary information shared among various agents and sensors. However, practical applications often encounter information disturbance which includes perception packet loss and time delays,…

Cited by 3SourcecodeScholar
2024

VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained Actions

NeurIPS 2024poster

Visual imitation learning (VIL) provides an efficient and intuitive strategy for robotic systems to acquire novel skills. Recent advancements in Vision Language Models (VLMs) have demonstrated remarkable performance in vision and language reasoning capabilities for VIL tasks. Despite the progress, c…

Cited by 5SourcePDFScholar
2023

Deep Interactive Full Transformer Framework for Point Cloud Registration

ICRA 2023poster

Point cloud registration is a crucial technology in the fields of robotics and computer vision. Despite the significant advances in point cloud registration enabled by Transformer-based methods, limitations persist due to indistinct feature extraction, noise sensitivity, and outlier handling. These…

Cited by 7SourcecodeScholar
2023

Multi-View Robust Collaborative Localization in High Outlier Ratio Scenes Based on Semantic Features

IROS 2023poster

Filtering out outlier data associations between local maps can improve the robustness and accuracy of multi-robot localization. When the overlap is low and the field of view difference is large, it is likely to produce outlier data associations between local maps, which will reduce the matching accu…

Cited by 4SourcecodeScholar
2023

PointGPT: Auto-regressively Generative Pre-training from Point Clouds

NeurIPS 2023poster

Large language models (LLMs) based on the generative pre-training transformer (GPT) have demonstrated remarkable effectiveness across a diverse range of downstream tasks. Inspired by the advancements of the GPT, we present PointGPT, a novel approach that extends the concept of GPT to point clouds, a…

2023

Rethinking Point Cloud Registration as Masking and Reconstruction

ICCV 2023poster

Point cloud registration is essential in computer vision and robotics. In this paper, a critical observation is made that the invisible parts of each point cloud can be directly utilized as inherent masks, and the aligned point cloud pair can be regarded as the reconstruction target. Motivated by th…

Cited by 13PDFcodeScholar
2023

SSGM: Spatial Semantic Graph Matching for Loop Closure Detection in Indoor Environments

IROS 2023poster

Capturing the semantics of objects and the topological relationship allows the robot to describe the scene more intelligently like a human and measure the similarity between scenes (loop closure detection) more accurately. However, many current semantic graph matching methods are based on walk descr…

Cited by 1SourcecodeScholar
2022

Aerial-Ground Robots Collaborative 3D Mapping in GNSS-Denied Environments

ICRA 2022poster

Collaborative heterogeneous robots are expected to perform comprehensive perception, mapping and coordination in search and rescue scenarios. The challenge of collaboration between heterogeneous robots lies in their huge differences in perception, mobility and processing capabilities. In this paper,…

Cited by 9SourceScholar
2022

HD-CCSOM: Hierarchical and Dense Collaborative Continuous Semantic Occupancy Mapping through Label Diffusion

IROS 2022poster

The collaborative operation of multiple robots can make up for the shortcomings of a single robot, such as limited field of perception or sensor failure. multirobots collaborative semantic mapping can enhance their comprehensive contextual understanding of the environment. However, existing multirob…

Cited by 16SourceScholar
2022

S-MKI: Incremental Dense Semantic Occupancy Reconstruction Through Multi-Entropy Kernel Inference

IROS 2022poster

Autonomous robots are often required to acquire high-level prior knowledge by continuously reconstructing the semantics and geometry of the surrounding scene, which is the basis of exploration and planning. Most existing continuous semantic mapping algorithms cannot distinguish potential differences…

Cited by 9SourceScholar
2022

Sem-Aug: Improving Camera-LiDAR Feature Fusion With Semantic Augmentation for 3D Vehicle Detection

RA-L 2022

Camera-LiDAR fusion provides precise distance measurements and fine-grained textures, making it a promising option for 3D vehicle detection in autonomous driving scenarios. Previous camera-LiDAR based 3D vehicle detection approaches mainly focused on employing image-based pre-trained models to fetch

Cited by 19SourceScholar
2021

MSTSL: Multi-Sensor Based Two-Step Localization in Geometrically Symmetric Environments

ICRA 2021poster

Symmetric environment is one of the most intractable and challenging scenarios for mobile robots to accomplish global localization tasks, due to the highly similar geometrical structures and insufficient distinctive features. Existing localization solutions in such scenarios either depend on pre-dep…

Cited by 21SourceScholar
2021

Robust Semantic Map Matching Algorithm Based on Probabilistic Registration Model

ICRA 2021poster

The matching and fusing of local maps generated by multiple robots can greatly enhance the performance of relative localization and collaborative mapping. Currently, existing semantic matching methods are partly based on classical iterative closet point (ICP), which typically fail in cases with larg…

Cited by 10SourceScholar
2021

Semantic Reinforced Attention Learning for Visual Place Recognition

ICRA 2021poster

Large-scale visual place recognition (VPR) is inherently challenging because not all visual cues in the image are beneficial to the task. In order to highlight the task-relevant visual cues in the feature embedding, the existing attention mechanisms are either based on artificial rules or trained in…

Cited by 70SourceScholar
2021

Tightly-Coupled Perception and Navigation of Heterogeneous Land-Air Robots in Complex Scenarios

ICRA 2021poster

In unstructured and unknown environments, heterogeneous robots must be able to perceive the environment, coordinate with each other and complete tasks collaboratively with onboard sensors. In this paper, a tightly-coupled perception and navigation framework is proposed for heterogeneous land-air rob…

Cited by 6SourceScholar
2020

A Hierarchical Framework for Collaborative Probabilistic Semantic Mapping

ICRA 2020poster

Performing collaborative semantic mapping is a critical challenge for cooperative robots to maintain a comprehensive contextual understanding of the surroundings. Most of the existing work either focus on single robot semantic mapping or collaborative geometry mapping. In this paper, a novel hierarc…

Cited by 38SourceScholar
2020

Collaborative Semantic Perception and Relative Localization Based on Map Matching

IROS 2020poster

In order to enable a team of robots to operate successfully, retrieving accurate relative transformation between robots is the fundamental requirement. So far, most research on relative localization mainly focus on geometry features such as points, lines and planes. To address this problem, collabor…

Cited by 22SourceScholar
2020

Day and Night Collaborative Dynamic Mapping in Unstructured Environment Based on Multimodal Sensors

ICRA 2020poster

Enabling long-term operation during day and night for collaborative robots requires a comprehensive understanding of the unstructured environment. Besides, in the dynamic environment, robots must be able to recognize dynamic objects and collaboratively build a global map. This paper proposes a novel…

Cited by 45SourceScholar