← Search

Saurabh Gupta

62 accepted papers

2025

AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time

EMNLP 2025

This paper presents AlphaOne ( 𝛼1 ), a universal framework for modulating reasoning progress in large reasoning models (LRMs) at test time. 𝛼1 first introduces 𝛼 moment, which represents the scaled thinking phase with a universal parameter 𝛼 .Within this scaled pre- 𝛼 moment phase, it dynamically sc

2025

Demonstrating MOSART: Opening Articulated Structures in the Real World

RSS 2025poster

What does it take to build mobile manipulation systems that can competently operate on previously unseen objects in previously unseen environments? This work answers this question using opening of articulated structures as a mobile manipulation testbed. Specifically, our focus is on the end-to-end p…

Cited by 0PDFcodeScholar
2025

How Do I Do That? Synthesizing 3D Hand Motion and Contacts for Everyday Interactions

CVPR 2025highlight

We tackle the novel problem of predicting 3D hand motion and contact maps (or Interaction Trajectories) given a single RGB view, action text, and a 3D contact point on the object as input. Our approach consists of (1) Interaction Codebook: a VQVAE model to learn a latent codebook of hand poses and c…

Cited by 2SourcePDFScholar
2025

KISS-SLAM: A Simple, Robust, and Accurate 3D LiDAR SLAM System With Enhanced Generalization Capabilities

IROS 2025

Robust and accurate localization and mapping of an environment using laser scanners, so-called LiDAR SLAM, is essential to many robotic applications. Early 3D LiDAR SLAM methods often exploited additional information from IMU or GNSS sensors to enhance localization accuracy and mitigate drift. Later

Cited by 17SourceScholar
2025

Learning Smooth Humanoid Locomotion through Lipschitz-Constrained Policies

IROS 2025

Reinforcement learning combined with sim-to-real transfer offers a general framework for developing locomotion controllers for legged robots. To facilitate successful deployment in the real world, smoothing techniques, such as low-pass filters and smoothness rewards, are often employed to develop po

Cited by 49SourcecodeScholar
2025

PhysGen3D: Crafting a Miniature Interactive World from a Single Image

CVPR 2025poster

Envisioning physically plausible outcomes from a single image requires a deep understanding of the world's dynamics. To address this, we introduce MiniTwin, a novel framework that transforms a single image into an amodal, camera-centric, interactive 3D scene. By combining advanced image-based geomet…

Cited by 3SourcePDFScholar
2025

Visual Sync: Multi‑Camera Synchronization via Cross‑View Object Motion

NeurIPS 2025poster

Today, people can easily record memorable moments, ranging from concerts, sports events, lectures, family gatherings, and birthday parties with multiple consumer cameras. However, synchronizing these cross‑camera streams remains challenging. Existing methods assume controlled settings, specific targ…

Cited by 0SourceScholar
2024

3D Reconstruction of Objects in Hands without Real World 3D Supervision

ECCV 2024poster

"Prior works for reconstructing hand-held objects from a single image train models on images paired with 3D shapes. Such data is challenging to gather in the real world at scale. Consequently, these approaches do not generalize well when presented with novel objects in in-the-wild settings. While 3D…

Cited by 2SourcePDFScholar
2024

Benchmarks and Challenges in Pose Estimation for Egocentric Hand Interactions with Objects

ECCV 2024poster

"We interact with the world with our hands and see it through our own (egocentric) perspective. A holistic understanding of such interactions from egocentric views is important for tasks in robotics, AR/VR, action recognition and motion generation. Accurately reconstructing such interactions in is c…

2024

Bootstrapping Autonomous Driving Radars with Self-Supervised Learning

CVPR 2024poster

The perception of autonomous vehicles using radars has attracted increased research interest due its ability to operate in fog and bad weather. However training radar models is hindered by the cost and difficulty of annotating large-scale radar data. To overcome this bottleneck we propose a self-sup…

2024

Diffusion Meets DAgger: Supercharging Eye-in-hand Imitation Learning

RSS 2024poster

A common failure mode for policies trained with imitation is compounding execution errors at test time. When the learned policy encounters states that are not present in the expert demonstrations, the policy fails, leading to degenerate behavior. The Dataset Aggregation, or DAgger approach to this p…

Cited by 15SourcePDFScholar
2024

Effectively Detecting Loop Closures using Point Cloud Density Maps

ICRA 2024poster

The ability to detect loop closures plays an essential role in any SLAM system. Loop closures allow correcting the drifting pose estimates from a sensor odometry pipeline. In this paper, we address the problem of effectively detecting loop closures in LiDAR SLAM systems in various environments with…

Cited by 15SourceScholar
2024

GOAT: GO to Any Thing

RSS 2024poster

In deployment scenarios such as homes and warehouses, mobile robots are expected to autonomously navigate for extended periods, seamlessly executing tasks articulated in terms that are intuitively understandable by human operators. We present GO To Any Thing (GOAT), a universal navigation system cap…

2024

Mitigating Perspective Distortion-induced Shape Ambiguity in Image Crops

ECCV 2024poster

"Objects undergo varying amounts of perspective distortion as they move across a camera’s field of view. Models for predicting 3D from a single image often work with crops around the object of interest and ignore the location of the object in the camera’s field of view. We note that ignoring this lo…

Cited by 3SourcePDFScholar
2024

Online Learning-Based Inertial Parameter Identification of Unknown Object for Model-Based Control of Wheeled Humanoids

RA-L 2024

Identifying the dynamic properties of manipulated objects is essential for safe and accurate robot control. Most methods rely on low-noise force-torque sensors, long exciting signals, and solving nonlinear optimization problems, making the estimation process slow. In this work, we propose a fast, on

Cited by 7SourceScholar
2024

PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation

ECCV 2024poster

"We present PhysGen, a novel image-to-video generation method that converts a single image and an input condition (, force and torque applied to an object in the image) to produce a realistic, physically plausible, and temporally consistent video. Our key insight is to integrate model-based physical…

2024

Toward Control of Wheeled Humanoid Robots with Unknown Payloads: Equilibrium Point Estimation via Real-to-Sim Adaptation

IROS 2024poster

Model-based controllers using a linearized model around the system’s equilibrium point is a common approach in the control of a wheeled humanoid due to their less computational load and ease of stability analysis. However, controlling a wheeled humanoid robot while it lifts an unknown object present…

Cited by 0SourceScholar
2023

Building Rearticulable Models for Arbitrary 3D Objects From 4D Point Clouds

CVPR 2023poster

We build rearticulable models for arbitrary everyday man-made objects containing an arbitrary number of parts that are connected together in arbitrary ways via 1-degree-of-freedom joints. Given point cloud videos of such everyday objects, our method identifies the distinct object parts, what parts a…

2023

ContactGen: Generative Contact Modeling for Grasp Generation

ICCV 2023poster

This paper presents a novel object-centric contact representation ContactGen for hand-object interaction. The ContactGen comprises 3 components: a contact map indicates the contact location, a part map represents the contact hand part, and a direction map tells the contact direction within each part…

Cited by 30PDFcodeScholar
2023

Exploiting Virtual Array Diversity for Accurate Radar Detection

ICASSP 2023accepted

Using millimeter-wave radars as a perception sensor provides self-driving cars with robust sensing capability in adverse weather. However, mmWave radars currently lack sufficient spatial resolution for semantic scene understanding. This paper introduces Radatron++, a system leverages cascaded MIMO (…

Cited by 0SourceScholar
2023

Look Ma, No Hands! Agent-Environment Factorization of Egocentric Videos

NeurIPS 2023poster

The analysis and use of egocentric videos for robotics tasks is made challenging by occlusion and the visual mismatch between the human hand and a robot end-effector. Past work views the human hand as a nuisance and removes it from the scene. However, the hand also provides a valuable signal for lea…

Cited by 24SourcePDFScholar
2022

CGF: Constrained Generation Framework for Query Rewriting in Conversational AI

EMNLP 2022industry

In conversational AI agents, Query Rewriting (QR) plays a crucial role in reducing user frictions and satisfying their daily demands. User frictions are caused by various reasons, such as errors in the conversational AI system, users’ accent or their abridged language. In this work, we present a nov…

2022

On-Device CPU Scheduling for Robot Systems

IROS 2022poster

Robots have to take highly responsive real-time actions, driven by complex decisions involving a pipeline of sensing, perception, planning, and reaction tasks. These tasks must be scheduled on resource-constrained devices such that the performance goals and the requirements of the application are me…

Cited by 5SourceScholar
2022

PAIGE: Personalized Adaptive Interactions Graph Encoder for Query Rewriting in Dialogue Systems

EMNLP 2022industry

Unexpected responses or repeated clarification questions from conversational agents detract from the users’ experience with technology meant to streamline their daily tasks. To reduce these frictions, Query Rewriting (QR) techniques replace transcripts of faulty queries with alternatives that lead t…

Cited by 1SourcePDFScholar
2022

Radatron: Accurate Detection Using Multi-Resolution Cascaded MIMO Radar

ECCV 2022poster

"Millimeter wave (mmWave) radars are becoming a more popular sensing modality in self-driving cars due to their favorable characteristics in adverse weather. Yet, they currently lack sufficient spatial resolution for semantic scene understanding. In this paper, we present Radatron, a system capable…

Cited by 30SourcePDFScholar
2022

TIDEE: Tidying Up Novel Rooms Using Visuo-Semantic Commonsense Priors

ECCV 2022poster

"We introduce TIDEE, an embodied agent that tidies up a disordered scene based on learned commonsense object placement and room arrangement priors. TIDEE explores a home environment, detects objects that are out of their natural place, infers plausible object contexts for them, localizes such contex…

2021

Contextual Rephrase Detection for Reducing Friction in Dialogue Systems

EMNLP 2021main

For voice assistants like Alexa, Google Assistant, and Siri, correctly interpreting users’ intentions is of utmost importance. However, users sometimes experience friction with these assistants, caused by errors from different system components or user errors such as slips of the tongue. Users tend…

2021

Graph Enhanced Query Rewriting for Spoken Language Understanding System

ICASSP 2021accepted

Query rewriting (QR) is an increasingly important component in voice assistant systems to reduce customer friction caused by errors in a spoken language understanding pipeline. These errors originate from various sources such as Automatic Speech Recognition (ASR) and Natural Language Understanding (…

Cited by 0SourceScholar
2021

Learned Visual Navigation for Under-Canopy Agricultural Robots

RSS 2021poster

This paper describes a system for visually guided autonomous navigation of under-canopy farm robots. Low-cost under-canopy robots can drive between crop rows under the plant canopy and accomplish tasks that are infeasible for over-the-canopy drones or larger agricultural equipment. However; autonomo…

Cited by 76SourcePDFScholar
2021

RB2: Robotic Manipulation Benchmarking with a Twist

NeurIPS 2021poster

Benchmarks offer a scientific way to compare algorithms using objective performance metrics. Good benchmarks have two features: (a) they should be widely useful for many research groups; (b) and they should produce reproducible findings. In robotic manipulation research, there is a trade-off between…

Cited by 25SourceScholar
2021

SEAL: Self-supervised Embodied Active Learning using Exploration and 3D Consistency

NeurIPS 2021poster

In this paper, we explore how we can build upon the data and models of Internet images and use them to adapt to robot vision without requiring any extra labels. We present a framework called Self-supervised Embodied Active Learning (SEAL). It utilizes perception models trained on internet images to…

Cited by 96SourcePDFScholar
2020

DeepRacer: Autonomous Racing Platform for Experimentation with Sim2Real Reinforcement Learning

ICRA 2020poster

DeepRacer is a platform for end-to-end experimentation with RL and can be used to systematically investigate the key challenges in developing intelligent control systems. Using the platform, we demonstrate how a 1/18th scale car can learn to drive autonomously using RL with a monocular camera. It is…

Cited by 76SourceScholar
2020

Efficient Bimanual Manipulation Using Learned Task Schemas

ICRA 2020poster

We address the problem of effectively composing skills to solve sparse-reward tasks in the real world. Given a set of parameterized skills (such as exerting a force or doing a top grasp at a location), our goal is to learn policies that invoke these skills to efficiently solve such tasks. Our insigh…

Cited by 84SourceScholar
2020

Intrinsic Motivation for Encouraging Synergistic Behavior

ICLR 2020poster

We study the role of intrinsic motivation as an exploration bias for reinforcement learning in sparse-reward synergistic tasks, which are tasks where multiple agents must work together to achieve a goal they could not individually. Our key idea is that a good guiding principle for intrinsic motivati…

Cited by 32SourceScholar
2020

Learning To Explore Using Active Neural SLAM

ICLR 2020poster

This work presents a modular and hierarchical approach to learn policies for exploring 3D environments, called `Active Neural SLAM'. Our approach leverages the strengths of both classical and learning-based methods, by using analytical path planners with learned SLAM module, and global and local pol…

Cited by 649SourcecodeScholar
2020

Through Fog High-Resolution Imaging Using Millimeter Wave Radar

CVPR 2020poster

This paper demonstrates high-resolution imaging using millimeter Wave (mmWave) radars that can function even in dense fog. We leverage the fact that mmWave signals have favorable propagation characteristics in low visibility conditions, unlike optical sensors like cameras and LiDARs which cannot pen…

Cited by 163PDFScholar
2020

Use the Force, Luke! Learning to Predict Physical Forces by Simulating Effects

CVPR 2020oral

When we humans look at a video of human-object interaction, we can not only infer what is happening but we can even extract actionable information and imitate those interactions. On the other hand, current recognition or geometric approaches lack the physicality of action representation. In this pap…

Cited by 57PDFcodeScholar
2019

Combining Optimal Control and Learning for Visual Navigation in Novel Environments

CoRL 2019

Model-based control is a popular paradigm for robot navigation because it can leverage a known dynamics model to efficiently plan robust robot trajectories. However, it is challenging to use model-based methods in settings where the environment is a priori unknown and can only be observed partially

Cited by 0SourcePDFScholar
2019

Segmenting Unknown 3D Objects from Real Depth Images using Mask R-CNN Trained on Synthetic Data

ICRA 2019poster

The ability to segment unknown objects in depth images has potential to enhance robot skills in grasping and object tracking. Recent computer vision research has demonstrated that Mask R-CNN can be trained to segment specific categories of objects in RGB images when massive hand-labeled datasets are…

Cited by 233SourcecodeScholar
2018

Factoring Shape, Pose, and Layout From the 2D Image of a 3D Scene

CVPR 2018poster

The goal of this paper is to take a single 2D image of a scene and recover the 3D structure in terms of a small set of factors: a layout representing the enclosing surfaces as well as a set of objects represented in terms of shape and pose. We propose a convolutional neural network-based approach to…

Cited by 157SourcePDFScholar
2018

Visual Memory for Robust Path Following

NeurIPS 2018oral

Humans routinely retrace a path in a novel environment both forwards and backwards despite uncertainty in their motion. In this paper, we present an approach for doing so. Given a demonstration of a path, a first network generates an abstraction of the path. Equipped with this abstraction, a second…

Cited by 64SourcePDFScholar
2017

Cognitive Mapping and Planning for Visual Navigation

CVPR 2017poster

We introduce a neural architecture for navigation in novel environments. Our proposed architecture learns to map from first-person views and plans a sequence of actions towards goals in the environment. The Cognitive Mapper and Planner (CMP) is based on two key ideas: a) a unified joint architecture…

Cited by 876PDFScholar
2015

Aligning 3D Models to RGB-D Images of Cluttered Scenes

CVPR 2015poster

The goal of this work is to represent objects in an RGB-D scene with corresponding 3D models from a library. We approach this problem by first detecting and segmenting object instances in the scene and then using a convolutional neural network (CNN) to predict the pose of the object. This CNN is tra…

Cited by 320SourcePDFScholar
2015

From Captions to Visual Concepts and Back

CVPR 2015poster

This paper presents a novel approach for automatically generating image descriptions: visual detectors, language models, and multimodal similarity models learnt directly from a dataset of image captions. We use multiple instance learning to train visual detectors for words that commonly occur in cap…