← Search

Martin Büchner

12 accepted papers

2026

Hoi! - A Multimodal Dataset for Force-Grounded, Cross-View Articulated Manipulation

CVPR 2026

We present a dataset for force-grounded, cross-view articulated manipulation that couples what is seen with what is done and what is felt during real human interaction. The dataset contains 3048 sequences across 381 articulated objects in 38 environments. Each object is operated in four embodiments

Cited by 0SourceScholar
2025

Articulated Object Estimation in the Wild

CoRL 2025poster

Understanding the 3D motion of articulated objects is essential in robotic scene understanding, mobile manipulation, and motion planning. Prior methods for articulation estimation have primarily focused on controlled settings, assuming either fixed camera viewpoints or direct observations of various…

Cited by 0SourceScholar
2025

Collaborative Dynamic 3D Scene Graphs for Open-Vocabulary Urban Scene Understanding

IROS 2025

Mapping and scene representation are fundamental to reliable planning and navigation in mobile robots. While purely geometric maps using voxel grids allow for general navigation, obtaining up-to-date spatial and semantically rich representations that scale to dynamic large-scale environments remains

Cited by 6SourceScholar
2025

MORE: Mobile Manipulation Rearrangement Through Grounded Language Reasoning

IROS 2025

Autonomous long-horizon mobile manipulation encompasses a multitude of challenges, including scene dynamics, unexplored areas, and error recovery. Recent works have leveraged foundation models for scene-level robotic reasoning and planning. However, the performance of these methods degrades when dea

Cited by 8SourceScholar
2025

OpenLex3D: A Tiered Benchmark for Open-Vocabulary 3D Scene Representations

NeurIPS 2025poster

3D scene understanding has been transformed by open-vocabulary language models that enable interaction via natural language. However, at present the evaluation of these representations is limited to datasets with closed-set semantics that do not capture the richness of language. This work presents O…

Cited by 0SourcecodeScholar
2025

Visual Loop Closure Detection Through Deep Graph Consensus

IROS 2025

Visual loop closure detection traditionally relies on place recognition methods to retrieve candidate loops that are validated using computationally expensive RANSAC-based geometric verification. As false positive loop closures significantly degrade downstream pose graph estimates, verifying a large

Cited by 1SourceScholar
2024

Collaborative Dynamic 3D Scene Graphs for Automated Driving

ICRA 2024poster

Maps have played an indispensable role in enabling safe and automated driving. Although there have been many advances on different fronts ranging from SLAM to semantics, building an actionable hierarchical semantic representation of urban dynamic scenes and processing information from multiple agent…

Cited by 27SourcecodeScholar
2024

Compositional Servoing by Recombining Demonstrations

ICRA 2024poster

Learning-based manipulation policies from image inputs often show weak task transfer capabilities. In contrast, visual servoing methods allow efficient task transfer in high-precision scenarios while requiring only a few demonstrations. In this work, we present a framework that formulates the visual…

Cited by 0SourceScholar
2024

Hierarchical Open-Vocabulary 3D Scene Graphs for Language-Grounded Robot Navigation

RSS 2024poster

Recent open-vocabulary robot mapping methods enrich dense geometric maps with pre-trained visual-language features. While these maps allow for the prediction of point-wise saliency maps when queried for a certain language concept, large-scale environments and abstract queries beyond the object level…

2024

Language-Grounded Dynamic Scene Graphs for Interactive Object Search With Mobile Manipulation

RA-L 2024

To fully leverage the capabilities of mobile manipulation robots, it is imperative that they are able to autonomously execute long-horizon tasks in large unexplored environments. While large language models (LLMs) have shown emergent reasoning skills on arbitrary tasks, existing work primarily conce

Cited by 100SourcecodeScholar
2023

Learning and Aggregating Lane Graphs for Urban Automated Driving

CVPR 2023poster

Lane graph estimation is an essential and highly challenging task in automated driving and HD map learning. Existing methods using either onboard or aerial imagery struggle with complex lane topologies, out-of-distribution scenarios, or significant occlusions in the image space. Moreover, merging ov…

Cited by 29SourcePDFScholar