← Search

Arsalan Mousavian

35 accepted papers

2026

PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation

CVPR 2026

Humans anticipate, from a glance and a contemplated action of their bodies, how the 3D world will respond, a capability that is equally vital for robotic manipulation. We introduce PointWorld, a large pre-trained 3D world model that unifies state and action in a shared 3D space as 3D point flows: gi

Cited by 0SourcecodeScholar
2025

Dexplore: Scalable Neural Control for Dexterous Manipulation from Reference Scoped Exploration

CoRL 2025poster

Hand–object motion-capture (MoCap) repositories provide abundant, contact-rich human demonstrations for scaling dexterous manipulation on robots. Yet demonstration inaccuracy and embodiment gaps between human and robot hands challenge direct policy learning. Existing pipelines adapt a three-stage wo…

Cited by 0SourceScholar
2025

VT-Refine: Learning Bimanual Assembly with Visuo-Tactile Feedback via Simulation Fine-Tuning

CoRL 2025poster

Humans excel at bimanual assembly tasks by adapting to rich tactile feedback—a capability that remains difficult to replicate in robots through behavioral cloning alone, due to the suboptimality and limited diversity of human demonstrations. In this work, we present VT-Refine, a visuo-tactile policy…

Cited by 0SourcecodeScholar
2024

DiffusionSeeder: Seeding Motion Optimization with Diffusion for Rapid Motion Planning

CoRL 2024poster

Running optimization across many parallel seeds leveraging GPU compute [2] have relaxed the need for a good initialization, but this can fail if the problem is highly non-convex as all seeds could get stuck in local minima. One such setting is collision-free motion optimization for robot manipulatio…

Cited by 4SourceScholar
2024

RoboPoint: A Vision-Language Model for Spatial Affordance Prediction in Robotics

CoRL 2024poster

From rearranging objects on a table to putting groceries into shelves, robots must plan precise action points to perform tasks accurately and reliably. In spite of the recent adoption of vision language models (VLMs) to control robot behavior, VLMs struggle to precisely articulate robot actions usin…

Cited by 53SourcecodeScholar
2023

CabiNet: Scaling Neural Collision Detection for Object Rearrangement with Procedural Scene Generation

ICRA 2023poster

We address the important problem of generalizing robotic rearrangement to clutter without any explicit object models. We first generate over 650K cluttered scenes-orders of magnitude more than prior work-in diverse everyday environments, such as cabinets and shelves. We render synthetic partial poin…

Cited by 27SourcecodeScholar
2023

Constrained Generative Sampling of 6-DoF Grasps

IROS 2023poster

Most state-of-the-art data-driven grasp sampling methods propose stable and collision-free grasps uniformly on the target object. For bin-picking, executing any of those reachable grasps is sufficient. However, for completing specific tasks, such as squeezing out liquid from a bottle, we want the gr…

Cited by 9SourcecodeScholar
2023

M2T2: Multi-Task Masked Transformer for Object-centric Pick and Place

CoRL 2023poster

With the advent of large language models and large-scale robotic datasets, there has been tremendous progress in high-level decision-making for object manipulation. These generic models are able to interpret complex tasks using language commands, but they often have difficulties generalizing to out-…

Cited by 23SourcecodeScholar
2023

ProgPrompt: Generating Situated Robot Task Plans using Large Language Models

ICRA 2023poster

Task planning can require defining myriad domain knowledge about the world in which a robot needs to act. To ameliorate that effort, large language models (LLMs) can be used to score potential next actions during task planning, and even generate action sequences directly, given an instruction in nat…

Cited by 893SourcecodeScholar
2022

IFOR: Iterative Flow Minimization for Robotic Object Rearrangement

CVPR 2022poster

Accurate object rearrangement from vision is a crucial problem for a wide variety of real-world robotics applications in unstructured environments. We propose IFOR, Iterative Flow Minimization for Robotic Object Rearrangement, an end-to-end method for the challenging problem of object rearrangement…

Cited by 59PDFcodeScholar
2022

Learning Robust Real-World Dexterous Grasping Policies via Implicit Shape Augmentation

CoRL 2022poster

Dexterous robotic hands have the capability to interact with a wide variety of household objects. However, learning robust real world grasping policies for arbitrary objects has proven challenging due to the difficulty of generating high quality training data. In this work, we propose a learning sys…

Cited by 32SourceScholar
2022

MegaPose: 6D Pose Estimation of Novel Objects via Render & Compare

CoRL 2022poster

We introduce MegaPose, a method to estimate the 6D pose of novel objects, that is, objects unseen during training. At inference time, the method only assumes knowledge of (i) a region of interest displaying the object in the image and (ii) a CAD model of the observed object. The contributions of thi…

Cited by 157SourcecodeScholar
2021

Contact-GraspNet: Efficient 6-DoF Grasp Generation in Cluttered Scenes

ICRA 2021poster

Grasping unseen objects in unconstrained, cluttered environments is an essential skill for autonomous robotic manipulation. Despite recent progress in full 6-DoF grasp learning, existing approaches often consist of complex sequential pipelines that possess several potential failure points and run-ti…

Cited by 424SourcecodeScholar
2021

Goal-Auxiliary Actor-Critic for 6D Robotic Grasping with Point Clouds

CoRL 2021poster

6D robotic grasping beyond top-down bin-picking scenarios is a challenging task. Previous solutions based on 6D grasp synthesis with robot motion planning usually operate in an open-loop setting, which are sensitive to grasp synthesis errors. In this work, we propose a new method for learning closed…

Cited by 55SourcecodeScholar
2021

NeRP: Neural Rearrangement Planning for Unknown Objects

RSS 2021poster

Robots will be expected to manipulate a wide variety of objects in complex and arbitrary ways as they become more widely used in human environments. As such; the rearrangement of objects has been noted to be an important benchmark for AI capabilities in recent years. We propose NeRP (Neural Rearrang…

Cited by 84SourcePDFScholar
2021

Object Rearrangement Using Learned Implicit Collision Functions

ICRA 2021poster

Robotic object rearrangement combines the skills of picking and placing objects. When object models are unavailable, typical collision-checking models may be unable to predict collisions in partial point clouds with occlusions, making generation of collision-free grasping or placement trajectories c…

Cited by 97SourceScholar
2021

RGB-D Local Implicit Function for Depth Completion of Transparent Objects

CVPR 2021poster

Majority of the perception methods in robotics require depth information provided by RGB-D cameras. However, standard 3D sensors fail to capture depth of transparent objects due to refraction and absorption of light. In this paper, we introduce a new approach for depth completion of transparent obje…

Cited by 98PDFcodeScholar
2021

RICE: Refining Instance Masks in Cluttered Environments with Graph Neural Networks

CoRL 2021poster

Segmenting unseen object instances in cluttered environments is an important capability that robots need when functioning in unstructured environments. While previous methods have exhibited promising results, they still tend to provide incorrect results in highly cluttered scenes. We postulate that…

Cited by 23SourcecodeScholar
2021

Reactive Human-to-Robot Handovers of Arbitrary Objects

ICRA 2021poster

Human-robot object handovers have been an actively studied area of robotics over the past decade; however, very few techniques and systems have addressed the challenge of handing over diverse objects with arbitrary appearance, size, shape, and deformability. In this paper, we present a vision-based…

Cited by 94SourceScholar
2021

Reactive Long Horizon Task Execution via Visual Skill and Precondition Models

IROS 2021poster

Zero-shot execution of unseen robotic tasks is important to allowing robots to perform a wide variety of tasks in human environments, but collecting the amounts of data necessary to train end-to-end policies in the real-world is often infeasible. We describe an approach for sim-to-real training that…

Cited by 21SourceScholar
2021

STORM: An Integrated Framework for Fast Joint-Space Model-Predictive Control for Reactive Manipulation

CoRL 2021oral

Sampling-based model-predictive control (MPC) is a promising tool for feedback control of robots with complex, non-smooth dynamics, and cost functions. However, the computationally demanding nature of sampling-based MPC algorithms has been a key bottleneck in their application to high-dimensional ro…

Cited by 152SourcecodeScholar
2021

Sim-to-Real for Robotic Tactile Sensing via Physics-Based Simulation and Learned Latent Projections

ICRA 2021poster

Tactile sensing is critical for robotic grasping and manipulation of objects under visual occlusion. However, in contrast to simulations of robot arms and cameras, current simulations of tactile sensors have limited accuracy, speed, and utility. In this work, we develop an efficient 3D finite elemen…

Cited by 69SourceScholar
2020

6-DOF Grasping for Target-driven Object Manipulation in Clutter

ICRA 2020poster

Grasping in cluttered environments is a fundamental but challenging robotic skill. It requires both reasoning about unseen object parts and potential collisions with the manipulator. Most existing data-driven approaches avoid this problem by limiting themselves to top-down planar grasps which is ins…

Cited by 263SourcecodeScholar
2020

Interpreting and Predicting Tactile Signals via a Physics-Based and Data-Driven Framework

RSS 2020poster

High-density afferents in the human hand have long been regarded as essential for human grasping and manipulation abilities. In contrast, robotic tactile sensors are typically used to provide low-density contact data, such as center-of-pressure and resultant force. Although useful, this data does no…

Cited by 29SourcePDFScholar
2020

LatentFusion: End-to-End Differentiable Reconstruction and Rendering for Unseen Object Pose Estimation

CVPR 2020poster

Current 6D object pose estimation methods usually require a 3D model for each object. These methods also require additional training in order to incorporate new objects. As a result, they are difficult to scale to a large number of objects and cannot be directly applied to unseen objects. We propose…

Cited by 170PDFScholar
2020

Learning RGB-D Feature Embeddings for Unseen Object Instance Segmentation

CoRL 2020

Segmenting unseen objects in cluttered scenes is an important skill that robots need to acquire in order to perform tasks in new environments. In this work, we propose a new method for unseen object instance segmentation by learning RGB-D feature embeddings from synthetic data. A metric learning los

Cited by 115SourcePDFScholar
2020

Self-supervised 6D Object Pose Estimation for Robot Manipulation

ICRA 2020poster

To teach robots skills, it is crucial to obtain data with supervision. Since annotating real world data is time-consuming and expensive, enabling robots to learn in a self- supervised way is important. In this work, we introduce a robot system for self-supervised 6D object pose estimation. Starting…

Cited by 239SourceScholar
2019

PoseRBPF: A Rao-Blackwellized Particle Filter for6D Object Pose Estimation

RSS 2019poster

Tracking 6D poses of objects from videos provides rich information to a robot in performing different tasks such as manipulation and navigation. In this work, we formulate the 6D object pose tracking problem in the Rao-Blackwellizedparticle filtering framework, where the 3D rotation and the 3D trans…

Cited by 0SourcePDFScholar
2019

The Best of Both Modes: Separately Leveraging RGB and Depth for Unseen Object Instance Segmentation

CoRL 2019

In order to function in unstructured environments, robots need the ability to recognize unseen novel objects. We take a step in this direction by tackling the problem of segmenting unseen object instances in tabletop environments. However, the type of large-scale real-world dataset required for this

Cited by 0SourcePDFScholar
2019

Visual Representations for Semantic Target Driven Navigation

ICRA 2019poster

What is a good visual representation for navigation? We study this question in the context of semantic visual navigation, which is the problem of a robot finding its way through a previously unseen environment to a target object, e.g. go to the refrigerator. Instead of acquiring a metric semantic ma…

Cited by 256SourcecodeScholar
2017

3D Bounding Box Estimation Using Deep Learning and Geometry

CVPR 2017poster

We present a method for 3D object detection and pose estimation from a single image. In contrast to current techniques that only regress the 3D orientation of an object, our method first regresses relatively stable 3D object properties using a deep convolutional neural network and then combines thes…

Cited by 1362PDFScholar
2017

Synthesizing Training Data for Object Detection in Indoor Scenes

RSS 2017poster

Detection of objects in cluttered indoor environments is one of the key enabling functionalities for service robots. The best performing object detection approaches in computer vision exploit deep Convolutional Neural Networks (CNN) to simultaneously detect and categorize the objects of interest in…

Cited by 285SourcePDFScholar