← Search

Baoru Huang

20 accepted papers

2026

AffordMatcher: Affordance Learning in 3D Scenes from Visual Signifiers

CVPR 2026

Affordance learning is a complex challenge in many applications, where existing approaches primarily focus on the geometric structures, visual knowledge, and affordance labels of objects to determine interactable regions. However, extending this learning capability to a scene is significantly more c

Cited by 0SourcecodeScholar
2026

SIGMA: A Physics-Based Benchmark for Gas Chimney Understanding in Seismic Images

CVPR 2026

Seismic images reconstruct subsurface reflectivity from field recordings, guiding exploration and reservoir monitoring. Gas chimneys are vertical anomalies caused by subsurface fluid migration. Understanding these phenomena is crucial for assessing hydrocarbon potential and avoiding drilling hazards

Cited by 0SourcecodeScholar
2026

StereoMamba: Real-Time and Robust Intraoperative Stereo Disparity Estimation Via Long-Range Spatial Dependencies

ICRA 2026poster

Stereo disparity estimation is crucial for obtaining depth information in robot-assisted minimally invasive surgery (RAMIS). While current deep learning methods have made significant advancements, challenges remain in achieving an optimal balance between accuracy, robustness, and inference speed. To…

2026

SurgCUT3R: Surgical Scene-Aware Continuous Understanding of Temporal 3D Representation

ICRA 2026poster

The reconstruction of surgical scenes from monocular endoscopic video is crucial for advancing robotic-assisted surgery, but applying state-of-the-art general-purpose reconstruction models is hindered by a severe lack of supervised training data and performance degradation over long sequences. To ad…

2025

EgoMusic-driven Human Dance Motion Estimation with Skeleton Mamba

ICCV 2025poster

Estimating human dance motion is a challenging task with various industrial applications. Recently, many efforts have focused on predicting human dance motion using either egocentric video or music as input. However, the task of jointly estimating human motion from both egocentric video and music re…

Cited by 0SourcePDFScholar
2025

FedEFM: Federated Endovascular Foundation Model with Unseen Data

ICRA 2025

In endovascular surgery, the precise identification of catheters and guidewires in X-ray images is essential for reducing intervention risks. However, accurately segmenting catheter and guidewire structures is challenging due to the limited availability of labeled data. Foundation models offer a pro

Cited by 3SourceScholar
2025

GraspMAS: Zero-Shot Language-driven Grasp Detection with Multi-Agent System

IROS 2025

Language-driven grasp detection has the potential to revolutionize human-robot interaction by allowing robots to understand and execute grasping tasks based on natural language commands. However, existing approaches face two key challenges. First, they often struggle to interpret complex text instru

Cited by 0SourcecodeScholar
2025

Hybrid Deep Reinforcement Learning for Radio Tracer Localisation in Robotic-Assisted Radioguided Surgery

ICRA 2025

Radioguided surgery, such as sentinel lymph node biopsy, relies on the precise localization of radioactive targets by non-imaging gamma/beta detectors. Manual radioactive target detection based on visual display or audible indication of gamma level is highly dependent on the ability of the surgeon t

Cited by 1SourceScholar
2025

Robotic-CLIP: Fine-Tuning CLIP on Action Data for Robotic Applications

ICRA 2025

Vision language models have played a key role in extracting meaningful features for various robotic applications. Among these, Contrastive Language-Image Pretraining (CLIP) is widely used in robotic tasks that require both vision and natural language understanding. However, CLIP was trained solely o

Cited by 11SourceScholar
2025

SplineFormer: An Explainable Transformer Network for Autonomous Endovascular Navigation

IROS 2025

Robot-assisted endovascular navigation provides significant advantages, including reduced radiation exposure for surgeons and improved patient safety. However, a major challenge is to control curvilinear instruments like guidewires precisely for smooth and accurate navigation while adapting to anato

Cited by 0SourceScholar
2025

Tracking Everything in Robotic-Assisted Surgery

ICRA 2025

Accurate tracking of tissues and instruments in videos is crucial for Robotic-Assisted Minimally Invasive Surgery (RAMIS), as it enables the robot to comprehend the surgical scene with precise locations and interactions of tissues and tools. Traditional keypoint-based sparse tracking is limited by f

Cited by 5SourcecodeScholar
2024

HabiCrowd: A High Performance Simulator for Crowd-Aware Visual Navigation

IROS 2024poster

Visual navigation, a foundational aspect of Embodied AI (E-AI) and robotics has been extensively studied in the past few years. While many 3D simulators have been introduced for the visual navigation tasks, scarcely works have combined human dynamics, creating the gap between simulation and real-wor…

Cited by 3SourcecodeScholar
2024

Language-Conditioned Affordance-Pose Detection in 3D Point Clouds

ICRA 2024poster

Affordance detection and pose estimation are of great importance in many robotic applications. Their combination helps the robot gain an enhanced manipulation capability, in which the generated pose can facilitate the corresponding affordance task. Previous methods for affodance-pose joint learning…

Cited by 17SourcecodeScholar
2024

Language-Driven 6-DoF Grasp Detection Using Negative Prompt Guidance

ECCV 2024oral

"6-DoF grasp detection has been a fundamental and challenging problem in robotic vision. While previous works have focused on ensuring grasp stability, they often do not consider human intention conveyed through natural language, hindering effective collaboration between robots and users in complex…

2024

Language-driven Grasp Detection with Mask-guided Attention

IROS 2024poster

Grasp detection is an essential task in robotics with various industrial applications. However, traditional methods often struggle with occlusions and do not utilize language for grasping. Incorporating natural language into grasp detection remains a challenging task and largely unexplored. To addre…

Cited by 1SourceScholar
2024

Lightweight Language-driven Grasp Detection using Conditional Consistency Model

IROS 2024

Language-driven grasp detection is a fundamental yet challenging task in robotics with various industrial applications. This work presents a new approach for language-driven grasp detection that leverages lightweight diffusion models to achieve fast inference time. By integrating diffusion processes

Cited by 12SourceScholar
2024

Open-Vocabulary Affordance Detection using Knowledge Distillation and Text-Point Correlation

ICRA 2024poster

Affordance detection presents intricate challenges and has a wide range of robotic applications. Previous works have faced limitations such as the complexities of 3D object shapes, the wide range of potential affordances on real-world objects, and the lack of open-vocabulary support for affordance u…

Cited by 10SourcecodeScholar
2023

Language-driven Scene Synthesis using Multi-conditional Diffusion Model

NeurIPS 2023poster

Scene synthesis is a challenging problem with several industrial applications. Recently, substantial efforts have been directed to synthesize the scene using human motions, room layouts, or spatial graphs as the input. However, few studies have addressed this problem from multiple modalities, especi…

2019

A Self-Adaptive Motion Scaling Framework for Surgical Robot Remote Control

RA-L 2019

Master-slave control is a common form of human-robot interaction for robotic surgery. To ensure seamless and intuitive control, a mechanism of self-adaptive motion scaling during teleoperaton is proposed in this letter. The operator can retain precise control when conducting delicate or complex mani

Cited by 44SourceScholar