← Search

Karthik Desingh

15 accepted papers

2025

AugInsert: Learning Robust Visual-Force Policies via Data Augmentation for Object Assembly Tasks

IROS 2025

Operating in unstructured environments like households requires robotic policies that are robust to out-of-distribution conditions. Although much work has been done in evaluating robustness for visuomotor policies, the robustness evaluation of a multisensory approach that includes force-torque sensi

Cited by 1SourcecodeScholar
2025

SuperQ-GRASP: Superquadrics-Based Grasp Pose Estimation on Larger Objects for Mobile-Manipulation

ICRA 2025

Grasp planning and estimation have been a longstanding research problem in robotics, with two main approaches to find graspable poses on the objects: 1) geometric approach, which relies on 3D models of objects and the gripper to estimate valid grasp poses, and 2) data-driven, learning-based approach

Cited by 3SourcecodeScholar
2024

Evaluating Robustness of Visual Representations for Object Assembly Task Requiring Spatio-Geometrical Reasoning

ICRA 2024poster

This paper primarily focuses on evaluating and benchmarking the robustness of visual representations in the context of object assembly tasks. Specifically, it investigates the alignment and insertion of objects with geometrical extrusions, commonly referred to as a peg-in-hole task. The accuracy req…

Cited by 2SourceScholar
2024

SlotGNN: Unsupervised Discovery of Multi-Object Representations and Visual Dynamics

ICRA 2024poster

Learning multi-object dynamics from visual data using unsupervised techniques is challenging due to the need for robust, object representations that can be learned through robot interactions. This paper presents a novel framework with two new architectures: SlotTransport for discovering object repre…

Cited by 3SourceScholar
2022

Break and Make: Interactive Structural Understanding Using LEGO Bricks

ECCV 2022poster

"Visual understanding of geometric structures with complex spatial relationships is a fundamental component of human intelligence. As children, we learn how to reason about structure not only from observation, but also by interacting with the world around us - by taking things apart and putting them…

2021

SORNet: Spatial Object-Centric Representations for Sequential Manipulation

CoRL 2021oral

Sequential manipulation tasks require a robot to perceive the state of an environment and plan a sequence of actions leading to a desired goal state, where the ability to reason about spatial relationships among object entities from raw sensor inputs is crucial. Prior works relying on explicit state…

Cited by 86SourcecodeScholar
2020

Parts-Based Articulated Object Localization in Clutter Using Belief Propagation

IROS 2020poster

Robots working in human environments must be able to perceive and act on challenging objects with articulations, such as a pile of tools. Articulated objects increase the dimensionality of the pose estimation problem, and partial observations under clutter create additional challenges. To address th…

Cited by 24SourceScholar
2019

Factored Pose Estimation of Articulated Objects using Efficient Nonparametric Belief Propagation

ICRA 2019poster

Robots working in human environments often encounter a wide range of articulated objects, such as tools, cabinets, and other jointed objects. Such articulated objects can take an infinite number of possible poses, as a point in a potentially high-dimensional continuous space. A robot must perceive t…

Cited by 48SourceScholar
2018

Gemsketch: Interactive Image-Guided Geometry Extraction from Point Clouds

ICRA 2018poster

We introduce an interactive system for extracting the geometries of generalized cylinders and cuboids from single-or multiple-view point clouds. Our proposed method is intuitive and only requires the object's silhouettes to be traced by the user. Leveraging the user's perceptual understanding of wha…

Cited by 6SourceScholar
2018

Semantic Mapping with Simultaneous Object Detection and Localization

IROS 2018poster

We present a filtering-based method for semantic mapping to simultaneously detect objects and localize their 6 degree-of-freedom pose. For our method, called Contextual Temporal Mapping (or CT-Map), we represent the semantic map as a belief over object classes and poses across an observed scene. Inf…

Cited by 39SourceScholar
2015

Axiomatic particle filtering for goal-directed robotic manipulation

IROS 2015poster

Manipulation tasks involving sequential pick-and-place actions in human environments remains an open problem for robotics. Central to this problem is the inability for robots to perceive in cluttered environments, where objects are physically touching, stacked, or occluded from the view. Such physic…

Cited by 36SourceScholar