← Search

K Madhava Krishna

58 accepted papers

2026

MonoMPC: Monocular Vision Based Navigation With Learned Collision Model and Risk-Aware Model Predictive Control

RA-L 2026

Navigating unknown environments with a single RGB camera is challenging, as the lack of depth information prevents reliable collision-checking. While some methods use estimated depth to build collision maps, we found that depth estimates from vision foundation models are too noisy for zero-shot navi

Cited by 1SourceScholar
2025

CrowdSurfer: Sampling Optimization Augmented with Vector-Quantized Variational AutoEncoder for Dense Crowd Navigation

ICRA 2025

Navigation amongst densely packed crowds remains a challenge for mobile robots. The complexity increases further if the environment layout changes, making the prior computed global plan infeasible. In this paper, we show that it is possible to dramatically enhance crowd navigation by just improving

Cited by 1SourcecodeScholar
2025

DG16M: A Large-Scale Dataset for Dual-Arm Grasping with Force-Optimized Grasps

IROS 2025

Dual-arm robotic grasping is crucial for handling large objects that require stable and coordinated manipulation. While single-arm grasping has been extensively studied, datasets tailored for dual-arm settings remain scarce. We introduce a large-scale dataset of 16 million dual-arm grasps, evaluated

Cited by 4SourcecodeScholar
2025

Da-Vil: Adaptive Dual-Arm Manipulation with Reinforcement Learning and Variable Impedance Control

ICRA 2025

Dual-arm manipulation is an area of growing interest in the robotics community. Enabling robots to perform tasks that require the coordinated use of two arms, is essential for complex manipulation tasks such as handling large objects, assembling components, and performing human-like interactions. Ho

Cited by 13SourcecodeScholar
2024

ATPPNet: Attention based Temporal Point cloud Prediction Network

ICRA 2024poster

Point cloud prediction is an important yet challenging task in the field of autonomous driving. The goal is to predict future point cloud sequences that maintain object structures while accurately representing their temporal motion. These predicted point clouds help in other subsequent tasks like ob…

Cited by 4SourceScholar
2024

Bi-level Trajectory Optimization on Uneven Terrains with Differentiable Wheel-Terrain Interaction Model

IROS 2024poster

Navigation of wheeled vehicles on uneven terrain necessitates going beyond the 2D approaches for trajectory planning. Specifically, it is essential to incorporate the full 6dof variation of vehicle pose and its associated stability cost in the planning process. To this end, most recent works aim to…

Cited by 3SourceScholar
2024

Constrained 6-DoF Grasp Generation on Complex Shapes for Improved Dual-Arm Manipulation

IROS 2024poster

Efficiently generating grasp poses tailored to specific regions of an object is vital for various robotic manipulation tasks, especially in a dual-arm setup. This scenario presents a significant challenge due to the complex geometries involved, requiring a deep understanding of the local geometry to…

Cited by 6SourcecodeScholar
2024

DiffPrompter: Differentiable Implicit Visual Prompts for Semantic-Segmentation in Adverse Conditions

IROS 2024poster

Semantic segmentation in adverse weather scenarios is a critical task for autonomous driving systems. While foundation models have shown promise, the need for specialized adaptors becomes evident for handling more challenging scenarios. We introduce DiffPrompter, a novel differentiable visual and la…

Cited by 0SourcecodeScholar
2024

Imagine2Servo: Intelligent Visual Servoing with Diffusion-Driven Goal Generation for Robotic Tasks

IROS 2024

Visual servoing, the method of controlling robot motion through feedback from visual sensors, has seen significant advancements with the integration of optical flow-based methods. However, its application remains limited by inherent challenges, such as the necessity for a target image at test time,

Cited by 3SourcecodeScholar
2024

LeGo-Drive: Language-enhanced Goal-oriented Closed-Loop End-to-End Autonomous Driving

IROS 2024poster

Existing Vision-Language Models (VLMs) produce long-term trajectory waypoints or directly control actions based on their perception input and language prompt. However, these VLMs are not explicitly aware of the constraints imposed by the scene or kinematics of the vehicle. As a result, the generated…

Cited by 3SourceScholar
2024

Leveraging Cycle-Consistent Anchor Points for Self-Supervised RGB-D Registration

ICRA 2024poster

With the rise in consumer depth cameras, a wealth of unlabeled RGB-D data has become available. This prompts the question of how to utilize this data for geometric reasoning of scenes. While many RGB-D registration methods rely on geometric and feature-based similarity, we take a different approach.…

Cited by 0SourceScholar
2024

Talk2BEV: Language-enhanced Bird’s-eye View Maps for Autonomous Driving

ICRA 2024poster

This work introduces Talk2BEV, a large vision-language model (LVLM)1 interface for bird’s-eye view (BEV) maps commonly used in autonomous driving. While existing perception systems for autonomous driving scenarios have largely focused on a pre-defined (closed) set of object categories and driving sc…

Cited by 77SourcecodeScholar
2023

GDIP: Gated Differentiable Image Processing for Object Detection in Adverse Conditions

ICRA 2023poster

Detecting objects under adverse weather and lighting conditions is crucial for the safe and continuous operation of an autonomous vehicle, and remains an unsolved problem. We present a Gated Differentiable Image Processing (GDIP) block, a domain-agnostic network architecture, which can be plugged in…

Cited by 68SourcecodeScholar
2023

Ground then Navigate: Language-guided Navigation in Dynamic Scenes

ICRA 2023poster

We investigate the Vision-and-Language Navigation (VLN) problem in the context of autonomous driving in outdoor settings. We solve the problem by explicitly grounding the navigable regions corresponding to the textual command. At each timestamp, the model predicts a segmentation mask corresponding t…

Cited by 27SourcecodeScholar
2022

CCO-VOXEL: Chance Constrained Optimization over Uncertain Voxel-Grid Representation for Safe Trajectory Planning

ICRA 2022poster

We present CCO-VOXEL: the very first chance-constrained optimization (CCO) algorithm that can compute trajectory plans with probabilistic safety guarantees in real-time directly on the voxel-grid representation of the world. CCO-VOXEL maps the distribution over the distance to the closest obstacle t…

Cited by 8SourcecodeScholar
2022

Drift Reduced Navigation with Deep Explainable Features

IROS 2022poster

Modern autonomous vehicles (AVs) often rely on vision, LIDAR, and even radar-based simultaneous localization and mapping (SLAM) frameworks for precise localization and navigation. However, modern SLAM frameworks often lead to unacceptably high levels of drift (i.e., localization error) when AVs obse…

Cited by 2SourcecodeScholar
2022

IndoLayout: Leveraging Attention for Extended Indoor Layout Estimation from an RGB Image

IROS 2022poster

In this work, we propose IndoLayout, a novel real-time approach for generating high-quality occupancy maps from an RGB image for indoor scenes. Such occupancy maps are often crucial for path-planning and mapping in indoor environments but are often built using only information contained in the ego v…

Cited by 0SourcecodeScholar
2022

Multi-Modal Model Predictive Control Through Batch Non-Holonomic Trajectory Optimization: Application to Highway Driving

RA-L 2022

Standard Model Predictive Control (MPC) or trajectory optimization approaches perform only a local search to solve a complex non-convex optimization problem. As a result, they cannot capture the multi-modal characteristic of human driving. A global optimizer can be a potential solution but is comput

Cited by 34SourcecodeScholar
2021

DRACO: Weakly Supervised Dense Reconstruction And Canonicalization of Objects

ICRA 2021poster

We present DRACO, a method for Dense Reconstruction And Canonicalization of Object shape from one or more RGB images. Canonical shape reconstruction— estimating 3D object shape in a coordinate space canonicalized for scale, rotation, and translation parameters—is an emerging paradigm that holds prom…

Cited by 6SourcecodeScholar
2021

Grounding Linguistic Commands to Navigable Regions

IROS 2021poster

Humans have a natural ability to effortlessly comprehend linguistic commands such as “park next to the yellow sedan” and instinctively know which region of the road the vehicle should navigate. Extending this ability to autonomous vehicles is the next step towards creating fully autonomous agents th…

Cited by 14SourcecodeScholar
2021

Modular Pipe Climber III with Three-Output Open Differential

IROS 2021poster

The paper introduces the novel Modular Pipe Climber III with a Three-Output Open Differential (3-OOD) mechanism to eliminate slipping of the tracks due to the changing cross-sections of the pipe. This will be achieved in any orientation of the robot. Previous pipe climbers use three-wheel/track modu…

Cited by 8SourceScholar
2021

RP-VIO: Robust Plane-based Visual-Inertial Odometry for Dynamic Environments

IROS 2021poster

Modern visual-inertial navigation systems (VINS) are faced with a critical challenge in real-world deployment: they need to operate reliably and robustly in highly dynamic environments. Current best solutions merely filter dynamic objects as outliers based on the semantics of the object category. Su…

Cited by 31SourcecodeScholar
2021

RTVS: A Lightweight Differentiable MPC Framework for Real-Time Visual Servoing

IROS 2021poster

Recent data-driven approaches to visual servoing have shown improved performances over classical methods due to precise feature matching and depth estimation. Some recent servoing approaches use a model predictive control (MPC) framework which generalise well to novel environments and are capable of…

Cited by 6SourceScholar
2021

RoRD: Rotation-Robust Descriptors and Orthographic Views for Local Feature Matching

IROS 2021poster

The use of local detectors and descriptors in typical computer vision pipelines works well until variations in viewpoint and appearance change become extreme. Past research in this area has typically focused on one of two approaches to this challenge: the use of projections into spaces more suitable…

Cited by 35SourcecodeScholar
2020

AutoLay: Benchmarking amodal layout estimation for autonomous driving

IROS 2020poster

Given an image or a video captured from a monocular camera, amodal layout estimation is the task of predicting semantics and occupancy in bird's eye view. The term amodal implies we also reason about entities in the scene that are occluded or truncated in image space. While several recent efforts ha…

Cited by 2SourceScholar
2020

Bi-Convex Approximation of Non-Holonomic Trajectory Optimization

ICRA 2020poster

Autonomous cars and fixed-wing aerial vehicles have the so-called non-holonomic kinematics which non-linearly maps control input to states. As a result, trajectory optimization with such a motion model becomes highly non-linear and non-convex. In this paper, we improve the computational tractability…

Cited by 10SourceScholar
2020

DFVS: Deep Flow Guided Scene Agnostic Image Based Visual Servoing

ICRA 2020poster

Existing deep learning based visual servoing approaches regress the relative camera pose between a pair of images. Therefore, they require a huge amount of training data and sometimes fine-tuning for adaptation to a novel scene. Furthermore, current approaches do not consider underlying geometry of…

Cited by 23SourceScholar
2020

LiDAR guided Small obstacle Segmentation

IROS 2020poster

Detecting small obstacles on the road is critical for autonomous driving. In this paper, we present a method to reliably detect such obstacles through a multi-modal framework of sparse LiDAR(VLP-16) and Monocular vision. LiDAR is employed to provide additional context in the form of confidence maps…

Cited by 33SourcecodeScholar
2020

Omnidirectional Tractable Three Module Robot

ICRA 2020poster

This paper introduces the Omnidirectional Tractable Three Module Robot for traversing inside complex pipe networks. The robot consists of three omnidirectional modules fixed 120° apart circumferentially which can rotate about their own axis allowing holonomic motion of the robot. The holonomic motio…

Cited by 18SourceScholar
2020

Reactive Navigation Under Non-Parametric Uncertainty Through Hilbert Space Embedding of Probabilistic Velocity Obstacles

RA-L 2020

The probabilistic velocity obstacle (PVO) extends the concept of velocity obstacle (VO) to work in uncertain dynamic environments. In this paper, we show how a robust model predictive control (MPC) with PVO constraints under non-parametric uncertainty can be made computationally tractable. At the co

Cited by 30SourceScholar
2020

Topological Mapping for Manhattan-like Repetitive Environments

ICRA 2020poster

We showcase a topological mapping framework for a challenging indoor warehouse setting. At the most abstract level, the warehouse is represented as a Topological Graph where the nodes of the graph represent a particular warehouse topological construct (e.g. rackspace, corridor) and the edges denote…

Cited by 12SourcecodeScholar
2020

Understanding Dynamic Scenes using Graph Convolution Networks

IROS 2020poster

We present a novel Multi-Relational Graph Convolutional Network (MRGCN) based framework to model on-road vehicle behaviors from a sequence of temporally ordered frames as grabbed by a moving monocular camera. The input to MRGCN is a multi-relational graph where the graph's nodes represent the active…

Cited by 34SourceScholar
2019

INFER: INtermediate representations for FuturE pRediction

IROS 2019poster

In urban driving scenarios, forecasting future trajectories of surrounding vehicles is of paramount importance. While several approaches for the problem have been proposed, the best-performing ones tend to require extremely detailed input representations (e.g. image sequences). As a result, such met…

Cited by 57SourceScholar
2019

Talk to the Vehicle: Language Conditioned Autonomous Navigation of Self Driving Cars

IROS 2019poster

We propose a novel pipeline that blends encodings from natural language and 3D semantic maps obtained from visual imagery to generate local trajectories that are executed by a low-level controller. The pipeline precludes the need for a prior registered map through a local waypoint generator neural n…

Cited by 30SourceScholar
2018

Beyond Pixels: Leveraging Geometry and Shape Cues for Online Multi-Object Tracking

ICRA 2018poster

This paper introduces geometry and object shape and pose costs for multi-object tracking in urban driving scenarios. Using images from a monocular camera alone, we devise pairwise costs for object tracks, based on several 3D cues such as object pose, shape, and motion. The proposed costs are agnosti…

Cited by 212SourcecodeScholar
2018

CalibNet: Geometrically Supervised Extrinsic Calibration using 3D Spatial Transformer Networks

IROS 2018poster

3D LiDARs and 2D cameras are increasingly being used alongside each other in sensor rigs for perception tasks. Before these sensors can be used to gather meaningful data, however, their extrinsics (and intrinsics) need to be accurately calibrated, as the performance of the sensor rig is extremely se…

Cited by 231SourcecodeScholar
2018

Constructing Category-Specific Models for Monocular Object-SLAM

ICRA 2018poster

We present a new paradigm for real-time object-oriented SLAM with a monocular camera. Contrary to previous approaches, that rely on object-level models, we construct category-level models from CAD collections which are now widely available. To alleviate the need for huge amounts of labeled data, we…

Cited by 66SourceScholar
2018

MergeNet: A Deep Net Architecture for Small Obstacle Discovery

ICRA 2018poster

We present here, a novel network architecture called MergeNet for discovering small obstacles for on-road scenes in the context of autonomous driving. The basis of the architecture rests on the central consideration of training with less amount of data since the physical setup and the annotation pro…

Cited by 29SourceScholar
2018

The Earth Ain't Flat: Monocular Reconstruction of Vehicles on Steep and Graded Roads from a Moving Camera

IROS 2018poster

Accurate localization of other traffic participants is a vital task in autonomous driving systems. State-of-the-art systems employ a combination of sensing modalities such as RGB cameras and LiDARs for localizing traffic participants, but monocular localization demonstrations have been confined to p…

Cited by 39SourceScholar
2018

Towards View-Invariant Intersection Recognition from Videos using Deep Network Ensembles

IROS 2018poster

This paper strives to answer the following question: Is it possible to recognize an intersection when seen from different road segments that constitute the intersection? An intersection or a junction typically is a meeting point of three or four road segments. Its recognition from a road segment tha…

Cited by 17SourceScholar
2017

Detachable modular robot capable of cooperative climbing and multi agent exploration

ICRA 2017poster

At the cross section of the fields of Uneven Terrain Navigation and Multi Agent Systems (MAS), in this work, a Detachable Compliant Modular Robot (DCMR) which can perform concurrent scene exploration by detaching into numerous parts, while preserving its ability to climb stairs is proposed and built…

Cited by 8SourceScholar
2017

Detecting, localizing, and recognizing trees with a monocular MAV: Towards preventing deforestation

ICRA 2017poster

We propose a novel pipeline for detecting, localizing, and recognizing trees with a quadcoptor equipped with monocular camera. The quadcoptor flies in an area of semidense plantation filled with many trees of more than 5 meter in height. Trees are detected on a per frame basis using state of the art…

Cited by 17SourceScholar
2017

Exploring convolutional networks for end-to-end visual servoing

ICRA 2017poster

Present image based visual servoing approaches rely on extracting hand crafted visual features from an image. Choosing the right set of features is important as it directly affects the performance of any approach. Motivated by recent breakthroughs in performance of data driven methods on recognition…

Cited by 104SourceScholar
2017

Multi-trajectory pose correspondences using scale-dependent topological analysis of pose-graphs

IROS 2017poster

This paper considers the problem of finding pose matches between trajectories of multiple robots in their respective coordinate frames or equivalent matches between trajectories obtained during different sessions. Pose correspondences between trajectories are mediated by common landmarks represented…

Cited by 1SourceScholar
2017

PRVO: Probabilistic Reciprocal Velocity Obstacle for multi robot navigation under uncertainty

IROS 2017poster

We present PRVO, a probabilistic variant of Reciprocal Velocity Obstacle (RVO) for decentralized multi-robot navigation under uncertainty. PRVO characterizes the space of velocities that would allow each robot to fulfill its share in collision avoidance with a specified probability. PRVO is modeled…

Cited by 77SourceScholar
2017

Pose induction for visual servoing to a novel object instance

IROS 2017poster

Present visual servoing approaches are instance specific i.e. they control camera motion between two views of the same object. However, in practical scenarios where a robot is required to handle various instances of a category, classical visual servoing techniques are less suitable. We formulate acr…

Cited by 9SourceScholar
2017

Reconstructing vehicles from a single image: Shape priors for road scene understanding

ICRA 2017poster

We present an approach for reconstructing vehicles from a single (RGB) image, in the context of autonomous driving. Though the problem appears to be ill-posed, we demonstrate that prior knowledge about how 3D shapes of vehicles project to an image can be used to reason about the reverse process, i.e…

Cited by 85SourceScholar
2017

Shape priors for real-time monocular object localization in dynamic environments

IROS 2017poster

Reconstruction of dynamic objects in a scene is a highly challenging problem in the context of SLAM. In this paper, we present a real-time monocular object localization system that estimates the shape and pose of dynamic objects in real-time, using video frames captured from a moving monocular camer…

Cited by 33SourceScholar
2016

Monocular reconstruction of vehicles: Combining SLAM with shape priors

ICRA 2016

Reasoning about objects in images and videos using 3D representations is re-emerging as a popular paradigm in computer vision. Specifically, in the context of scene understanding for roads, 3D vehicle detection and tracking from monocular videos still needs a lot of attention to enable practical app

Cited by 50SourceScholar
2016

Plantation monitoring and yield estimation using autonomous quadcopter for precision agriculture

ICRA 2016

Recently, quadcopters with their advance sensors and imaging capabilities have become an imperative part of the precision agriculture. In this work, we have described a framework which performs plantation monitoring and yield estimation using the supervised learning approach, while autonomously navi

Cited by 44SourceScholar
2016

Rolling shutter and motion blur removal for depth cameras

ICRA 2016

Structured light range sensors (SLRS) like the Microsoft Kinect have electronic rolling shutters (ERS). The output of such a sensor while in motion is subject to significant motion blur (MB) and rolling shutter (RS) distortion. Most robotic literature still does not explicitly model this distortion,

Cited by 10SourceScholar
2015

Autonomous navigation of generic monocular quadcopter in natural environment

ICRA 2015poster

Autonomous navigation of generic monocular quadcopter in the natural environment requires sophisticated mechanism for perception, planning and control. In this work, we have described a framework which performs perception using monocular camera and generates minimum time collision free trajectory an…

Cited by 45SourceScholar