← Search

Shubham Sonawani

7 accepted papers

2024

Diff-Control: A Stateful Diffusion-based Policy for Imitation Learning

IROS 2024poster

While imitation learning provides a simple and effective framework for policy learning, acquiring consistent action during robot execution remains a challenging task. Existing approaches primarily focus on either modifying the action representation at data curation stage or altering the model itself…

Cited by 2SourcecodeScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration0

ICRA 2024poster

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man…

Cited by 259SourcecodeScholar
2024

SiSCo: Signal Synthesis for Effective Human-Robot Communication Via Large Language Models

IROS 2024poster

Effective human-robot collaboration hinges on robust communication channels, with visual signaling playing a pivotal role due to its intuitive appeal. Yet, the creation of visually intuitive cues often demands extensive resources and specialized knowledge. The emergence of Large Language Models (LLM…

Cited by 1SourceScholar
2024

iRoCo: Intuitive Robot Control From Anywhere Using a Smartwatch

ICRA 2024poster

This paper introduces iRoCo (intuitive Robot Control) – a framework for ubiquitous human-robot collaboration using a single smartwatch and smartphone. By integrating probabilistic differentiable filters, iRoCo optimizes a combination of precise robot control and unrestricted user movement from ubiqu…

Cited by 2SourcecodeScholar
2023

Anytime, Anywhere: Human Arm Pose from Smartwatch Data for Ubiquitous Robot Control and Teleoperation

IROS 2023poster

This work devises an optimized machine learning approach for human arm pose estimation from a single smart-watch. Our approach results in a distribution of possible wrist and elbow positions, which allows for a measure of uncertainty and the detection of multiple possible arm posture solutions, i.e.…

Cited by 6SourceScholar
2023

Projecting Robot Intentions Through Visual Cues: Static vs. Dynamic Signaling

IROS 2023poster

Augmented and mixed-reality techniques harbor a great potential for improving human-robot collaboration. Visual signals and cues may be projected to a human partner in order to explicitly communicate robot intentions and goals. However, it is unclear what type of signals support such a process and w…

Cited by 5SourceScholar
2022

Modularity through Attention: Efficient Training and Transfer of Language-Conditioned Policies for Robot Manipulation

CoRL 2022poster

Language-conditioned policies allow robots to interpret and execute human instructions. Learning such policies requires a substantial investment with regards to time and compute resources. Still, the resulting controllers are highly device-specific and cannot easily be transferred to a robot with di…

Cited by 25SourcecodeScholar