← Search

Binbin Xu

17 accepted papers

2026

HIPPo: Harnessing Image-To-3D Priors for Model-Free Zero-Shot 6D Pose Estimation

ICRA 2026poster

This work focuses on the problem of 6D pose estimation for novel objects when a reference 3D model or posed reference images are not available. While existing methods can estimate the precise 6D pose of objects, they heavily rely on curated CAD models or reference images, the preparation of which is…

2025

HIPPo: Harnessing Image-to-3D Priors for Model-Free Zero-Shot 6D Pose Estimation

RA-L 2025

This work focuses on the problem of 6D pose estimation for novel objects when a reference 3D model or posed reference images are not available. While existing methods can estimate the precise 6D pose of objects, they heavily rely on curated CAD models or reference images, the preparation of which is

Cited by 4SourceScholar
2025

UnPose: Uncertainty-Guided Diffusion Priors for Zero-Shot Pose Estimation

CoRL 2025poster

Estimating the 6D pose of novel objects is a fundamental yet challenging problem in robotics, often relying on access to object CAD models. However, acquiring such models can be costly and impractical. Recent approaches aim to bypass this requirement by leveraging strong priors from founda…

Cited by 0SourceScholar
2024

FuncGrasp: Learning Object-Centric Neural Grasp Functions from Single Annotated Example Object

ICRA 2024poster

We present FuncGrasp, a framework that can infer dense yet reliable grasp configurations for unseen objects using one annotated object and single-view RGB-D observation via categorical priors. Unlike previous works that only transfer a set of grasp poses, FuncGrasp aims to transfer infinite configur…

Cited by 4SourceScholar
2024

Identifying Optimal Launch Sites of High-Altitude Latex-Balloons using Bayesian Optimisation for the Task of Station-Keeping

IROS 2024

Station-keeping tasks for high-altitude balloons show promise in areas such as ecological surveys, atmospheric analysis, and communication relays. However, identifying the optimal time and position to launch a latex high-altitude balloon is still a challenging and multifaceted problem. For example,

Cited by 0SourceScholar
2024

MOSE: Monocular Semantic Reconstruction Using NeRF-Lifted Noisy Priors

RA-L 2024

Accurately reconstructing dense and semantically annotated 3D meshes from monocular images remains a challenging task due to the lack of geometry guidance and imperfect view-dependent 2D priors. Though we have witnessed recent advancements in implicit neural scene representations enabling precise 2D

Cited by 0SourceScholar
2024

Toward General Object-level Mapping from Sparse Views with 3D Diffusion Priors

CoRL 2024poster

Object-level mapping builds a 3D map of objects in a scene with detailed shapes and poses from multi-view sensor observations. Conventional methods struggle to build complete shapes and estimate accurate poses due to partial occlusions and sensor noise. They require dense observations to cover all…

Cited by 3SourcecodeScholar
2023

Accurate and Interactive Visual-Inertial Sensor Calibration with Next-Best-View and Next-Best-Trajectory Suggestion

IROS 2023poster

Visual-Inertial (VI) sensors are popular in robotics, self-driving vehicles, and augmented and virtual reality applications. In order to use them for any computer vision or state-estimation task, a good calibration is essential. However, collecting informative calibration data in order to render the…

Cited by 3SourcecodeScholar
2023

Finding Things in the Unknown: Semantic Object-Centric Exploration with an MAV

ICRA 2023poster

Exploration of unknown space with an autonomous mobile robot is a well-studied problem. In this work we broaden the scope of exploration, moving beyond the pure geometric goal of uncovering as much free space as possible. We believe that for many practical applications, exploration should be context…

Cited by 23SourceScholar
2023

Incremental Dense Reconstruction From Monocular Video With Guided Sparse Feature Volume Fusion

RA-L 2023

Incrementally recovering 3D dense structures from monocular videos is of paramount importance since it enables various robotics and AR applications. Feature volumes have recently been shown to enable efficient and accurate incremental dense reconstruction without the need to first estimate depth, bu

Cited by 12SourceScholar
2022

Learning to Complete Object Shapes for Object-level Mapping in Dynamic Scenes

IROS 2022poster

In this paper, we propose a novel object-level mapping system that can simultaneously segment, track, and reconstruct objects in dynamic scenes. It can further predict and complete their full geometries by conditioning on reconstructions from depth inputs and a category-level shape prior with the ai…

Cited by 13SourceScholar
2022

Visual-Inertial Multi-Instance Dynamic SLAM with Object-level Relocalisation

IROS 2022poster

In this paper, we present a tightly-coupled visual-inertial object-level multi-instance dynamic SLAM system. Even in extremely dynamic scenes, it can robustly optimise for the camera pose, velocity, IMU biases and build a dense 3D reconstruction object-level map of the environment. Our system can ro…

Cited by 16SourcecodeScholar
2019

MID-Fusion: Octree-based Object-Level Multi-Instance Dynamic SLAM

ICRA 2019poster

We propose a new multi-instance dynamic RGB-D SLAM system using an object-level octree-based volumetric representation. It can provide robust camera tracking in dynamic environments and at the same time, continuously estimate geometric, semantic, and motion properties for arbitrary objects in the sc…

Cited by 240SourcecodeScholar
2017

Spatio-Temporal Video Completion in Spherical Image Sequences

RA-L 2017

Spherical cameras are widely used due to their full 360$^\circ$ fields of view. However, a common but severe problem is that anything carrying the camera is always included in the view, occluding visual information. In this letter, we propose a novel method to remove such occlusions in videos taken

Cited by 9SourceScholar