← Search

Roberto Martín-Martín

77 accepted papers

2026

CAVER: Curious AudioVisual Exploring Robot

ICRA 2026poster

Multimodal audiovisual perception can enable new avenues for robotic manipulation, from better material classification to the imitation of demonstrations for which only audio signals are available (e.g., playing a tune by ear). However, to unlock such multimodal potential, robots need to learn the c…

2026

CoDex: Learning Compositional Dexterous Functional Manipulation without Demonstrations

ICRA 2026poster

In this work, we study Compositional Dexterous Functional Object Manipulation (CD-FOM): tasks such as aiming and actuating a spray bottle on a plant or a glue gun on wood, which require both actuating an object's internal mechanism and controlling its pose to apply the object's function to the envir…

2026

DataMIL: Selecting Data for Robot Imitation Learning with Datamodels

ICLR 2026poster

Recently, the robotics community has amassed ever larger and more diverse datasets to train generalist policies. However, while these policies achieve strong mean performance across a variety of tasks, they often underperform on individual, specialized tasks and require further tuning on newly acqui…

Cited by 0SourceScholar
2026

Factored Latent Action World Models

ICML 2026poster

Learning latent actions from action-free video has emerged as a powerful paradigm for scaling up controllable world model learning. Latent actions provide a natural interface for users to iteratively generate and manipulate videos. However, most existing approaches rely on monolithic inverse and for…

Cited by 0SourceScholar
2026

Large-Language-Model-Guided State Estimation for Partially Observable Task and Motion Planning

ICRA 2026poster

Robot planning in partially observable environments, where not all objects are known or visible, is a challenging problem, as it requires reasoning under uncertainty through partially observable Markov decision processes. During the execution of a computed plan, a robot may unexpectedly observe task…

2026

Mash, Spread, Slice! Learning to Manipulate Object States Via Visual Spatial Progress

ICRA 2026poster

Most robot manipulation focuses on changing the kinematic state of objects: picking, placing, opening, or rotating them. However, a wide range of real-world manipulation tasks involve a different class of object state change—such as mashing, spreading, or slicing—where the object’s physical and visu…

2026

Memory Retrieval in Visuomotor Policies for Long-Horizon Robot Control

RSS 2026poster

General-purpose robots operating in partially observable environments such as homes require memory to support long-term autonomy. They must recall different types of past information, such as where objects were placed, which subtasks have already been completed by a human partner, and when an applia…

Cited by 0SourceScholar
2026

MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos

ICRA 2026poster

We aim to enable humanoid robots to efficiently solve new manipulation tasks from a few video examples. In-context learning (ICL) is a promising framework for achieving this goal due to its test-time data efficiency and rapid adaptability. However, current ICL methods rely on labor-intensive teleope…

2026

Mixed-Initiative Dialog for Human-Robot Collaborative Manipulation

ICRA 2026poster

Effective robotic systems for long-horizon human-robot collaboration must adapt to a wide range of human partners, whose physical behavior, willingness to assist, and understanding of the robot's capabilities may change over time. This demands a tightly coupled communication loop that grants both ag…

2026

MoMaGen: Generating Demonstrations under Soft and Hard Constraints for Multi-Step Bimanual Mobile Manipulation

ICLR 2026poster

Imitation learning from large-scale, diverse human demonstrations has been shown to be effective for training robots, but collecting such data is costly and time-consuming. This challenge intensifies for multi-step bimanual mobile manipulation, where humans must teleoperate both the mobile base and…

Cited by 0SourcecodeScholar
2026

OopsieVerse: A Safety Benchmark with Damage-Aware Simulation for Robot Manipulation

RSS 2026poster

While robotic manipulation capabilities have advanced rapidly, physical safety remains a major barrier to deploying household robots: task success is insufficient if the robot damages itself or its surroundings. Simulation offers a harm-free alternative to costly and dangerous real-world training an…

2026

Rapid Adaptation of Particle Dynamics for Generalized Deformable Object Mobile Manipulation

ICRA 2026poster

We address the challenge of learning to manipulate deformable objects with unknown dynamics. In non-rigid objects, the dynamics parameters define how they react to interactions --how they stretch, bend, compress, and move-- and they are critical to determining the optimal actions to perform a manipu…

2026

Searching in Space and Time: Unified Memory-Action Loops for Open-World Object Retrieval

ICRA 2026poster

Service robots must retrieve objects in dynamic, open-world settings where requests may reference attributes (“the red mug”), spatial context (“the mug on the table”), or past states (“the mug that was here yesterday”). Existing approaches capture only parts of this problem: scene graphs capture spa…

2025

BUMBLE: Unifying Reasoning and Acting with Vision-Language Models for Building-wide Mobile Manipulation

ICRA 2025

To operate at a building scale, service robots must perform long-horizon mobile manipulation tasks by navigating to different rooms, accessing multiple floors, and interacting with a wide and unseen range of everyday objects. We refer to these tasks as Building-wide Mobile Manipulation. To tackle th

Cited by 31SourceScholar
2025

CASPER: Inferring Diverse Intents for Assistive Teleoperation with Vision Language Models

CoRL 2025poster

Assistive teleoperation, where control is shared between a human and a robot, enables efficient and intuitive human-robot collaboration in diverse and unstructured environments. A central challenge in real-world assistive teleoperation is for the robot to infer a wide range of human intentions from…

Cited by 0SourcecodeScholar
2025

COLLAGE: Adaptive Fusion-based Retrieval for Augmented Policy Learning

CoRL 2025poster

In this work, we study the problem of data retrieval for few-shot imitation learning: select data from a large dataset to train a performant policy for a specific task, given only a few target demonstrations. Prior methods retrieve data using a single-feature distance heuristic, assuming that the be…

Cited by 0SourcecodeScholar
2025

Deep Reinforcement Learning for Robotics: A Survey of Real-World Successes

AAAI 2025technical

Reinforcement learning (RL), particularly its combination with deep neural networks referred to as deep RL (DRL), has shown tremendous promise across a wide range of applications, suggesting its potential for enabling the development of sophisticated robotic behaviors. Robotics problems, however, po…

Cited by 48SourcePDFScholar
2025

FLaRe: Achieving Masterful and Adaptive Robot Policies with Large-Scale Reinforcement Learning Fine-Tuning

ICRA 2025

In recent years, the Robotics field has initiated several efforts toward building generalist robot policies through large-scale multi-task Behavior Cloning. However, direct deployments of these policies have led to unsatisfactory performance, where the policy struggles with unseen states and tasks.

Cited by 57SourcecodeScholar
2025

GeT-USE: Learning Generalized Tool Usage for Bimanual Mobile Manipulation via Simulated Embodiment Extensions

IROS 2025

The ability to use random objects as tools in a generalizable manner is a missing piece in robots’ intelligence today to boost their versatility and problem-solving capabilities. State-of-the-art robotic tool usage methods focused on procedurally generating or crowd-sourcing datasets of tools for a

Cited by 0SourceScholar
2025

RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies

CoRL 2025oral

Comprehensive, unbiased, and comparable evaluation of modern generalist policies is uniquely challenging: existing approaches for robot benchmarking typically rely on heavy standardization, either by specifying fixed evaluation tasks and environments, or by hosting centralized "robot challenges", an…

Cited by 0SourceScholar
2025

SLAC: Simulation-Pretrained Latent Action Space for Whole-Body Real-World RL

CoRL 2025poster

Building capable household and industrial robots requires mastering the control of versatile, high-degree-of-freedom (DoF) systems such as mobile manipulators. While reinforcement learning (RL) holds promise for autonomously acquiring robot control policies, scaling it to high-DoF embodiments remain…

Cited by 0SourceScholar
2025

SafeMimic: Towards Safe and Autonomous Human-to-Robot Imitation for Mobile Manipulation

RSS 2025poster

For robots to become efficient helpers in the home, they must learn to perform new mobile manipulation tasks simply by watching humans perform them. Learning from a single video demonstration from a human is challenging as the robot needs to first extract from the demo what needs to be done and how,…

Cited by 0PDFScholar
2024

BEHAVIOR Vision Suite: Customizable Dataset Generation via Simulation

CVPR 2024highlight

The systematic evaluation and understanding of computer vision models under varying conditions require large amounts of data with comprehensive and customized labels which real-world vision datasets rarely satisfy. While current synthetic data generators offer a promising alternative particularly fo…

2024

BaRiFlex: A Robotic Gripper with Versatility and Collision Robustness for Robot Learning

IROS 2024poster

We present a new approach to robot hand design specifically suited to enable robot learning methods and daily tasks in human environments. We introduce BaRiFlex, an innovative gripper design that alleviates the issues caused by unexpected contact and collisions during robot learning, offering collis…

Cited by 3SourceScholar
2024

DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

RSS 2024poster

The creation of large, diverse, high-quality robot manipulation datasets is an important stepping stone on the path toward more capable and robust robotic manipulation policies. However, creating such datasets is challenging: collecting robot manipulation data in diverse environments poses logistica…

Cited by 216SourcePDFScholar
2024

Disentangled Unsupervised Skill Discovery for Efficient Hierarchical Reinforcement Learning

NeurIPS 2024poster

A hallmark of intelligent agents is the ability to learn reusable skills purely from unsupervised interaction with the environment. However, existing unsupervised skill discovery methods often learn entangled skills where one skill variable simultaneously influences many entities in the environment,…

Cited by 2SourcePDFScholar
2024

Hierarchical Point Attention for Indoor 3D Object Detection

ICRA 2024poster

3D object detection is an essential vision technique for various robotic systems, such as augmented reality and domestic robots. Transformers as versatile network architectures have recently seen great success in 3D point cloud object detection. However, the lack of hierarchy in a plain transformer…

Cited by 1SourceScholar
2024

Learning to Look: Seeking Information for Decision Making via Policy Factorization

CoRL 2024poster

Many robot manipulation tasks require active or interactive exploration behavior in order to be performed successfully. Such tasks are ubiquitous in embodied domains, where agents must actively search for the information necessary for each stage of a task, e.g., moving the head of the robot to find…

Cited by 0SourceScholar
2024

Model-Based Runtime Monitoring with Interactive Imitation Learning

ICRA 2024poster

Robot learning methods have recently made great strides, but generalization and robustness challenges still hinder their widespread deployment. Failing to detect and address potential failures renders state-of-the-art learning systems not combat-ready for high-stakes tasks. Recent advances in intera…

Cited by 20SourcecodeScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration

ICRA 2024

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man

Cited by 910SourcecodeScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration0

ICRA 2024poster

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man…

Cited by 259SourcecodeScholar
2024

ScrewMimic: Bimanual Imitation from Human Videos with Screw Space Projection

RSS 2024poster

Bimanual manipulation is a longstanding challenge in robotics due to the large number of degrees of freedom and the strict spatial and temporal synchronization required to generate meaningful behavior. Humans learn bimanual manipulation skills by watching other humans and by refining their abilities…

Cited by 17SourcePDFScholar
2024

SkiLD: Unsupervised Skill Discovery Guided by Factor Interactions

NeurIPS 2024poster

Unsupervised skill discovery carries the promise that an intelligent agent can learn reusable skills through autonomous, reward-free interactions with environments. Existing unsupervised skill discovery methods learn skills by encouraging distinguishable behaviors that cover diverse states. However,…

Cited by 1SourcePDFScholar
2024

ULIP-2: Towards Scalable Multimodal Pre-training for 3D Understanding

CVPR 2024poster

Recent advancements in multimodal pre-training have shown promising efficacy in 3D representation learning by aligning multimodal features across 3D shapes their 2D counterparts and language descriptions. However the methods used by existing frameworks to curate such multimodal data in particular la…

2023

M-EMBER: Tackling Long-Horizon Mobile Manipulation via Factorized Domain Transfer

ICRA 2023poster

In this paper, we propose a novel method to create visuomotor mobile manipulation solutions to long-horizon activities. We propose to leverage the recent advances in robot simulation to train robust visual solutions in simulation that can transfer to the real world. While previous works have shown s…

Cited by 14SourceScholar
2023

MUTEX: Learning Unified Policies from Multimodal Task Specifications

CoRL 2023poster

Humans use different modalities, such as speech, text, images, videos, etc., to communicate their intent and goals with teammates. For robots to become better assistants, we aim to endow them with the ability to follow instructions and understand tasks specified by their human partners. Most robotic…

Cited by 65SourcecodeScholar
2023

MaskViT: Masked Visual Pre-Training for Video Prediction

ICLR 2023poster

The ability to predict future visual observations conditioned on past observations and motor commands can enable embodied agents to plan solutions to a variety of tasks in complex environments. This work shows that we can create good video prediction models by pre-training transformers via masked vi…

Cited by 135SourcePDFScholar
2023

Modeling Dynamic Environments with Scene Graph Memory

ICML 2023poster

Embodied AI agents that search for objects in large environments such as households often need to make efficient decisions by predicting object locations based on partial information. We pose this as a new type of link prediction problem: link prediction on partially observable dynamic graphs Our gr…

Cited by 14SourcePDFScholar
2023

Procedure-Aware Pretraining for Instructional Video Understanding

CVPR 2023poster

Our goal is to learn a video representation that is useful for downstream procedure understanding tasks in instructional videos. Due to the small amount of available annotations, a key challenge in procedure understanding is to be able to extract from unlabeled videos the procedural knowledge such a…

2023

Task-Driven Graph Attention for Hierarchical Relational Object Navigation

ICRA 2023poster

Embodied AI agents in large scenes often need to navigate to find objects. In this work, we study a naturally emerging variant of the object navigation task, hierarchical relational object navigation (HRON), where the goal is to find objects specified by logical predicates organized in a hierarchica…

Cited by 7SourceScholar
2023

ULIP: Learning a Unified Representation of Language, Images, and Point Clouds for 3D Understanding

CVPR 2023poster

The recognition capabilities of current state-of-the-art 3D models are limited by datasets with a small number of annotated data and a pre-defined set of categories. In its 2D counterpart, recent advances have shown that similar problems can be significantly alleviated by employing knowledge from ot…

2022

A Dual Representation Framework for Robot Learning with Human Guidance

CoRL 2022poster

The ability to interactively learn skills from human guidance and adjust behavior according to human preference is crucial to accelerating robot learning. But human guidance is an expensive resource, calling for methods that can learn efficiently. In this work, we argue that learning is more efficie…

Cited by 15SourceScholar
2022

BEHAVIOR-1K: A Benchmark for Embodied AI with 1,000 Everyday Activities and Realistic Simulation

CoRL 2022oral

We present BEHAVIOR-1K, a comprehensive simulation benchmark for human-centered robotics. BEHAVIOR-1K includes two components, guided and motivated by the results of an extensive survey on "what do you want robots to do for you?". The first is the definition of 1,000 everyday activities, grounded in…

Cited by 205SourceScholar
2021

BEHAVIOR: Benchmark for Everyday Household Activities in Virtual, Interactive, and Ecological Environments

CoRL 2021poster

We introduce BEHAVIOR, a benchmark for embodied AI with 100 activities in simulation, spanning a range of everyday household chores such as cleaning, maintenance, and food preparation. These activities are designed to be realistic, diverse and complex, aiming to reproduce the challenges that agents…

Cited by 176SourceScholar
2021

Deep Affordance Foresight: Planning Through What Can Be Done in the Future

ICRA 2021poster

Planning in realistic environments requires searching in large planning spaces. Affordances are a powerful concept to simplify this search, because they model what actions can be successful in a given situation. However, the classical notion of affordance is not suitable for long horizon planning be…

Cited by 94SourcecodeScholar
2021

Differentiable Factor Graph Optimization for Learning Smoothers

IROS 2021poster

A recent line of work has shown that end-to-end optimization of Bayesian filters can be used to learn state estimators for systems whose underlying models are difficult to hand-design or tune, while retaining the core advantages of probabilistic state estimation. As an alternative approach for state…

Cited by 28SourceScholar
2021

Error-Aware Imitation Learning from Teleoperation Data for Mobile Manipulation

CoRL 2021poster

In mobile manipulation (MM), robots can both navigate within and interact with their environment and are thus able to complete many more tasks than robots only capable of navigation or manipulation. In this work, we explore how to apply imitation learning (IL) to learn continuous visuo-motor policie…

Cited by 66SourceScholar
2021

LASER: Learning a Latent Action Space for Efficient Reinforcement Learning

ICRA 2021poster

The process of learning a manipulation task depends strongly on the action space used for exploration: posed in the incorrect action space, solving a task with reinforcement learning can be drastically inefficient. Additionally, similar tasks or instances of the same task family impose latent manifo…

Cited by 69SourceScholar
2021

Learning Multi-Arm Manipulation Through Collaborative Teleoperation

ICRA 2021poster

Imitation Learning (IL) is a powerful paradigm to teach robots to perform manipulation tasks by allowing them to learn from human demonstrations collected via teleoperation, but has mostly been limited to single-arm manipulation. However, many real-world tasks require multiple arms, such as lifting…

Cited by 58SourceScholar
2021

Probabilistic Visual Navigation with Bidirectional Image Prediction

IROS 2021poster

Humans can robustly follow a visual trajectory defined by a sequence of images (i.e. a video) regardless of substantial changes in the environment or the presence of obstacles. We aim at endowing similar visual navigation capabilities to mobile robots solely equipped with a RGB fisheye camera. We pr…

Cited by 8SourceScholar
2021

ReLMoGen: Integrating Motion Generation in Reinforcement Learning for Mobile Manipulation

ICRA 2021poster

Many Reinforcement Learning (RL) approaches use joint control signals (positions, velocities, torques) as action space for continuous control tasks. We propose to lift the action space to a higher level in the form of subgoals for a motion generator (a combination of motion planner and trajectory ex…

Cited by 81SourceScholar
2021

Robot Navigation in Constrained Pedestrian Environments using Reinforcement Learning

ICRA 2021poster

Navigating fluently around pedestrians is a necessary capability for mobile robots deployed in human environments, such as buildings and homes. While research on social navigation has focused mainly on the scalability with the number of pedestrians in open spaces, typical indoor environments present…

Cited by 98SourceScholar
2021

Semantic and Geometric Modeling with Neural Message Passing in 3D Scene Graphs for Hierarchical Mechanical Search

ICRA 2021poster

Searching for objects in indoor organized environments such as homes or offices is part of our everyday activities. When looking for a desired object, we reason about the rooms and containers the object is likely to be in; the same type of container will have a different probability of containing th…

Cited by 37SourceScholar
2021

What Matters in Learning from Offline Human Demonstrations for Robot Manipulation

CoRL 2021oral

Imitating human demonstrations is a promising approach to endow robots with various manipulation capabilities. While recent advances have been made in imitation learning and batch (offline) reinforcement learning, a lack of open-source human datasets and reproducible learning methods make assessing…

Cited by 523SourcecodeScholar
2021

iGibson 1.0: A Simulation Environment for Interactive Tasks in Large Realistic Scenes

IROS 2021poster

We present iGibson 1.0, a novel simulation environment to develop robotic solutions for interactive tasks in large-scale realistic scenes. Our environment contains 15 fully interactive home-sized scenes with 108 rooms populated with rigid and articulated objects. The scenes are replicas of real-worl…

Cited by 193SourceScholar
2021

iGibson 2.0: Object-Centric Simulation for Robot Learning of Everyday Household Tasks

CoRL 2021poster

Recent research in embodied AI has been boosted by the use of simulation environments to develop and train robot learning approaches. However, the use of simulation has skewed the attention to tasks that only require what robotics simulators can simulate: motion and physical contact. We present iGib…

Cited by 268SourceScholar
2020

6-PACK: Category-level 6D Pose Tracker with Anchor-Based Keypoints

ICRA 2020poster

We present 6-PACK, a deep learning approach to category-level 6D object pose tracking on RGB-D data. Our method tracks in real time novel object instances of known object categories such as bowls, laptops, and mugs. 6-PACK learns to compactly represent an object by a handful of 3D keypoints, based o…

Cited by 190SourcecodeScholar
2020

GTI: Learning to Generalize across Long-Horizon Tasks from Human Demonstrations

RSS 2020poster

Imitation learning is an effective and safe technique to train robot policies in the real world because it does not depend on an expensive random exploration process. However, due to the lack of exploration, learning policies that generalize beyond the demonstrated behaviors is still an open challen…

Cited by 174SourcePDFScholar
2020

Interactive Gibson Benchmark: A Benchmark for Interactive Navigation in Cluttered Environments

RA-L 2020

We present Interactive Gibson Benchmark, the first comprehensive benchmark for training and evaluating Interactive Navigation solutions. Interactive Navigation tasks are robot navigation problems where physical interaction with objects (e.g., pushing) is allowed and even encouraged to reach the goal

Cited by 211SourceScholar
2020

JRMOT: A Real-Time 3D Multi-Object Tracker and a New Large-Scale Dataset

IROS 2020poster

Robots navigating autonomously need to perceive and track the motion of objects and other agents in its surroundings. This information enables planning and executing robust and safe trajectories. To facilitate these processes, the motion should be perceived in 3D Cartesian space. However, most recen…

Cited by 106SourcecodeScholar
2020

Multimodal Sensor Fusion with Differentiable Filters

IROS 2020poster

Leveraging multimodal information with recursive Bayesian filters improves performance and robustness of state estimation, as recursive filters can combine different modalities according to their uncertainties. Prior work has studied how to optimally fuse different sensor modalities with analytical…

Cited by 67SourceScholar
2020

Visuomotor Mechanical Search: Learning to Retrieve Target Objects in Clutter

IROS 2020poster

When searching for objects in cluttered environments, it is often necessary to perform complex interactions in order to move occluding objects out of the way and fully reveal the object of interest and make it graspable. Due to the complexity of the physics involved and the lack of accurate models o…

Cited by 51SourceScholar
2019

Deep Local Trajectory Replanning and Control for Robot Navigation

ICRA 2019poster

We present a navigation system that combines ideas from hierarchical planning and machine learning. The system uses a traditional global planner to compute optimal paths towards a goal, and a deep local trajectory planner and velocity controller to compute motion commands. The latter components of t…

Cited by 88SourceScholar
2019

Deep Visual MPC-Policy Learning for Navigation

RA-L 2019

Humans can routinely follow a trajectory defined by a list of images/landmarks. However, traditional robot navigation methods require accurate mapping of the environment, localization, and planning. Moreover, these methods are sensitive to subtle changes in the environment. In this letter, we propos

Cited by 114SourceScholar
2019

HRL4IN: Hierarchical Reinforcement Learning for Interactive Navigation with Mobile Manipulators

CoRL 2019

Most common navigation tasks in human environments require auxiliary arm interactions, e.g. opening doors, pressing buttons and pushing obstacles away. This type of navigation tasks, which we call Interactive Navigation, requires the use of mobile manipulators: mobile bases with manipulation capabil

Cited by 0SourcePDFScholar
2019

Mechanical Search: Multi-Step Retrieval of a Target Object Occluded by Clutter

ICRA 2019poster

When operating in unstructured environments such as warehouses, homes, and retail centers, robots are frequently required to interactively search for and retrieve specific objects from cluttered bins, shelves, or tables. Mechanical Search describes the class of tasks where the goal is to locate and…

Cited by 141SourceScholar
2019

Regression Planning Networks

NeurIPS 2019poster

Recent learning-to-plan methods have shown promising results on planning directly from observation space. Yet, their ability to plan for long-horizon tasks is limited by the accuracy of the prediction model. On the other hand, classical symbolic planners show remarkable capabilities in solving long-…

2019

Social-BiGAT: Multimodal Trajectory Forecasting using Bicycle-GAN and Graph Attention Networks

NeurIPS 2019poster

Predicting the future trajectories of multiple interacting pedestrians in a scene has become an increasingly important problem for many different applications ranging from control of autonomous vehicles and social robots to security and surveillance. This problem is compounded by the presence of soc…

Cited by 838SourcePDFScholar
2019

VUNet: Dynamic Scene View Synthesis for Traversability Estimation Using an RGB Camera

RA-L 2019

We present VUNet, a novel view(VU) synthesis method for mobile robots in dynamic environments, and its application to the estimation of future traversability. Our method predicts future images for given virtual robot velocity commands using only RGB images at previous and current time steps. The fut

Cited by 40SourceScholar
2019

Variable Impedance Control in End-Effector Space: An Action Space for Reinforcement Learning in Contact-Rich Tasks

IROS 2019poster

Reinforcement Learning (RL) of contact-rich manipulation tasks has yielded impressive results in recent years. While many studies in RL focus on varying the observation space or reward model, few efforts focused on the choice of action space (e.g. joint or end-effector space, position, velocity, etc…

Cited by 231SourcecodeScholar
2018

Physics-Based Selection of Informative Actions for Interactive Perception

ICRA 2018poster

Interactive perception exploits the correlation between forceful interactions and changes in the observed signals to extract task-relevant information from the sensor stream. Finding the most informative interactions to perceive complex objects, like articulated mechanisms, is challenging because th…

Cited by 9SourceScholar
2017

Cross-modal interpretation of multi-modal sensor streams in interactive perception based on coupled recursion

IROS 2017poster

We present an online system to perceive kinematic properties of articulated objects from multi-modal sensor streams. The novelty of our system is that it leverages multi-modal information in a cross-modal manner: instead of simply fusing information from different modalities, sensor streams are inte…

Cited by 16SourceScholar
2016

Probabilistic multi-class segmentation for the Amazon Picking Challenge

IROS 2016poster

We present a method for multi-class segmentation from RGB-D data in a realistic warehouse picking setting. The method computes pixel-wise probabilities and combines them to find a coherent object segmentation. It reliably segments objects in cluttered scenarios, even when objects are translucent, re…

Cited by 81SourceScholar