← Search

Abhinav Valada

86 accepted papers

2026

CoVAR: Co-Generation of Video and Action for Robotic Manipulation Via Multi-Modal Diffusion

ICRA 2026poster

We present a method to generate video–action pairs that follow text instructions, starting from an initial image observation and the robot’s joint states. Our approach automatically provides action labels for video diffusion mod- els, overcoming the common lack of action annotations and enabling the…

2026

DOS: Distilling Observable Softmaps of Zipfian Prototypes for Self-Supervised Point Representation

AAAI 2026technical

Recent advances in self-supervised learning (SSL) have shown tremendous potential for learning 3D point cloud representations without human annotations. However, SSL for 3D point clouds still faces critical challenges due to irregular geometry, shortcut-prone reconstruction, and unbalanced semantics

Cited by 0SourcePDFScholar
2026

Efficient Learning of Object Placement With Intra-Category Transfer

RA-L 2026

Efficient learning from demonstration for long horizon tasks remains an open challenge in robotics. While significant effort has been directed toward learning trajectories, a recent resurgence of object-centric approaches has demonstrated improved sample efficiency, enabling transferable robotic ski

Cited by 1SourceScholar
2026

Efficient Learning of Object Placement with Intra-Category Transfer

ICRA 2026poster

Efficient learning from demonstration for long-horizon tasks remains an open challenge in robotics. While significant effort has been directed toward learning trajectories, a recent resurgence of object-centric approaches has demonstrated improved sample efficiency, enabling transferable robotic ski…

2026

ForecastOcc: Vision-Based Semantic Occupancy Forecasting

ICRA 2026poster

Autonomous driving requires forecasting both geometry and semantics over time to effectively reason about future environment states. Existing vision-based occupancy forecasting methods focus on motion-related categories such as static and dynamic objects, while semantic information remains largely a…

2026

Hoi! - A Multimodal Dataset for Force-Grounded, Cross-View Articulated Manipulation

CVPR 2026

We present a dataset for force-grounded, cross-view articulated manipulation that couples what is seen with what is done and what is felt during real human interaction. The dataset contains 3048 sequences across 381 articulated objects in 38 environments. Each object is operated in four embodiments

Cited by 0SourceScholar
2026

MSG: Multi-Stream Generative Policies for Sample-Efficient Robotic Manipulation

RA-L 2026

Generative robot policies such as Flow Matching offer flexible, multi-modal policy learning but are sample-inefficient. Although object-centric policies improve sample efficiency, it does not resolve this limitation. In this work, we propose Multi-Stream Generative Policy (MSG), an inference-time co

Cited by 0SourceScholar
2026

ParkDiffusion++: Ego Intention Conditioned Joint Trajectory Prediction for Automated Parking Using Diffusion Models

ICRA 2026poster

Automated parking is a challenging operational domain for advanced driver assistance systems, requiring robust scene understanding and interaction reasoning. The key challenge is twofold: (i)predict multiple plausible ego intentions according to context and (ii)for each intention, predict the joint …

Cited by 0Scholar
2026

ROVER: A Multi-Season Dataset for Visual SLAM

ICRA 2026poster

Robust Simultaneous Localization and Mapping (SLAM) is a crucial enabler for autonomous navigation in natural, semi-structured environments such as parks and gardens. However, these environments present unique challenges for SLAM due to frequent seasonal changes, varying light conditions, and dense …

2026

Scaling Single Human Demonstrations for Imitation Learning Using Generative Foundational Models

ICRA 2026poster

Imitation learning is a popular paradigm to teach robots new tasks, but collecting robot demonstrations through teleoperation or kinesthetic teaching is tedious and time-consuming. In contrast, directly demonstrating a task using our human embodiment is much easier and data is available in abundance…

2026

UP-Fuse: Uncertainty-guided LiDAR-Camera Fusion for 3D Panoptic Segmentation

RSS 2026poster

LiDAR-camera fusion enhances 3D panoptic segmentation by leveraging camera images to complement sparse LiDAR scans, but it also introduces a critical failure mode. Under adverse conditions, degradation or failure of the camera sensor can significantly compromise the reliability of the perception sys…

Cited by 0SourceScholar
2025

A Good Foundation is Worth Many Labels: Label-Efficient Panoptic Segmentation

RA-L 2025

A key challenge for the widespread application of learning-based models for robotic perception is to significantly reduce the required amount of annotated training data while achieving accurate predictions. This is essential not only to decrease operating costs but also to speed up deployment time.

Cited by 8SourcecodeScholar
2025

Articulated Object Estimation in the Wild

CoRL 2025poster

Understanding the 3D motion of articulated objects is essential in robotic scene understanding, mobile manipulation, and motion planning. Prior methods for articulation estimation have primarily focused on controlled settings, assuming either fixed camera viewpoints or direct observations of various…

Cited by 0SourceScholar
2025

Collaborative Dynamic 3D Scene Graphs for Open-Vocabulary Urban Scene Understanding

IROS 2025

Mapping and scene representation are fundamental to reliable planning and navigation in mobile robots. While purely geometric maps using voxel grids allow for general navigation, obtaining up-to-date spatial and semantically rich representations that scale to dynamic large-scale environments remains

Cited by 6SourceScholar
2025

DiWA: Diffusion Policy Adaptation with World Models

CoRL 2025poster

Fine-tuning diffusion policies with reinforcement learning (RL) presents significant challenges. The long denoising sequence for each action prediction impedes effective reward propagation. Additionally, standard RL methods require millions of physical interaction steps, making fine-tuning even more…

Cited by 0SourceScholar
2025

Evidential Uncertainty Estimation for Multi-Modal Trajectory Prediction

IROS 2025

Accurate trajectory prediction is crucial for autonomous driving, yet uncertainty in agent behavior and perception noise makes it inherently challenging. While multi-modal trajectory prediction models generate multiple plausible future paths with associated probabilities, effectively quantifying unc

Cited by 5SourceScholar
2025

Label-Efficient LiDAR Panoptic Segmentation

IROS 2025

A main bottleneck of learning-based robotic scene understanding methods is the heavy reliance on extensive annotated training data, which often limits their generalization ability. In LiDAR panoptic segmentation, this challenge becomes even more pronounced due to the need to simultaneously address b

Cited by 1SourceScholar
2025

Label-Efficient LiDAR Semantic Segmentation with 2D-3D Vision Transformer Adapters

IROS 2025

LiDAR semantic segmentation models are typically trained from random initialization as universal pre-training is hindered by the lack of large, diverse datasets. Moreover, most point cloud segmentation architectures incorporate custom network layers, limiting the transferability of advances from vis

Cited by 7SourceScholar
2025

MORE: Mobile Manipulation Rearrangement Through Grounded Language Reasoning

IROS 2025

Autonomous long-horizon mobile manipulation encompasses a multitude of challenges, including scene dynamics, unexplored areas, and error recovery. Recent works have leveraged foundation models for scene-level robotic reasoning and planning. However, the performance of these methods degrades when dea

Cited by 8SourceScholar
2025

Motion Forecasting via Model-Based Risk Minimization

ICRA 2025

Forecasting the future trajectories of surrounding agents is crucial for autonomous vehicles to ensure safe, efficient, and comfortable route planning. While model ensembling has improved prediction accuracy in various fields, its application in trajectory prediction is limited due to the multi-moda

Cited by 5SourceScholar
2025

Multi-Scale Neighborhood Occupancy Masked Autoencoder for Self-Supervised Learning in LiDAR Point Clouds

CVPR 2025poster

Masked autoencoders (MAE) have shown tremendous potential for self-supervised learning (SSL) in vision and beyond. However, point clouds from LiDARs used in automated driving are particularly challenging for MAEs since large areas of the 3D volume are empty. Consequently, existing work suffers from…

Cited by 0SourcePDFScholar
2025

Open-Set LiDAR Panoptic Segmentation Guided by Uncertainty-Aware Learning

IROS 2025

Autonomous vehicles that navigate in open-world environments may encounter previously unseen object classes. However, most existing LiDAR panoptic segmentation models rely on closed-set assumptions, failing to detect unknown object instances. In this work, we propose ULOPS, an uncertainty-guided ope

Cited by 3SourceScholar
2025

OpenLex3D: A Tiered Benchmark for Open-Vocabulary 3D Scene Representations

NeurIPS 2025poster

3D scene understanding has been transformed by open-vocabulary language models that enable interaction via natural language. However, at present the evaluation of these representations is limited to datasets with closed-set semantics that do not capture the richness of language. This work presents O…

Cited by 0SourcecodeScholar
2025

ParkDiffusion: Heterogeneous Multi-Agent Multi-Modal Trajectory Prediction for Automated Parking using Diffusion Models

IROS 2025

Automated parking is a critical feature of Advanced Driver Assistance Systems (ADAS), where accurate trajectory prediction is essential to bridge perception and planning modules. Despite its significance, research in this domain remains relatively limited, with most existing studies concentrating on

Cited by 5SourceScholar
2025

PseudoTouch: Efficiently Imaging the Surface Feel of Objects for Robotic Manipulation

ICRA 2025

Tactile sensing is vital for human dexterous manipulation, however, it has not been widely used in robotics. Compact, low-cost sensing platforms can facilitate a change, but unlike their popular optical counterparts, they are difficult to deploy in high-fidelity tasks due to their low signal dimensi

Cited by 1SourceScholar
2025

RoboEnvision: A Long-Horizon Video Generation Model for Multi-Task Robot Manipulation

IROS 2025

We address the problem of generating long-horizon videos for robotic manipulation tasks. Text-to-video diffusion models have made significant progress in photorealism, language understanding, and motion generation but struggle with long-horizon robotic tasks. Recent works use video diffusion models

Cited by 11SourceScholar
2025

Robotic Task Ambiguity Resolution via Natural Language Interaction

IROS 2025

Language-Conditioned robotic policies allow users to specify tasks using natural language. While much research has focused on improving the action prediction of language-conditioned policies, reasoning about task descriptions has been largely overlooked. Ambiguous task descriptions often lead to dow

Cited by 5SourceScholar
2025

Stochasticity in Motion: An Information-Theoretic Approach to Trajectory Prediction

IROS 2025

In autonomous driving, accurate motion prediction is crucial for safe and efficient motion planning. To ensure safety, planners require reliable uncertainty estimates of the predicted behavior of surrounding agents, yet this aspect has received limited attention. In particular, decomposing uncertain

Cited by 5SourceScholar
2025

Taxonomy-Aware Continual Semantic Segmentation in Hyperbolic Spaces for Open-World Perception

RA-L 2025

Semantic segmentation models are typically trained on a fixed set of classes, limiting their applicability in open-world scenarios. Class-incremental semantic segmentation aims to update models with emerging new classes while preventing catastrophic forgetting of previously learned ones. However, ex

Cited by 6SourceScholar
2025

Visual Loop Closure Detection Through Deep Graph Consensus

IROS 2025

Visual loop closure detection traditionally relies on place recognition methods to retrieve candidate loops that are validated using computationally expensive RANSAC-based geometric verification. As false positive loop closures significantly degrade downstream pose graph estimates, verifying a large

Cited by 1SourceScholar
2025

Whole-Body Teleoperation for Mobile Manipulation at Zero Added Cost

RA-L 2025

Demonstration data plays a key role in learning complex behaviors and training robotic foundation models. While effective control interfaces exist for static manipulators, data collection remains cumbersome and time intensive for mobile manipulators due to their large number of degrees of freedom. W

Cited by 12SourceScholar
2024

A Point-Based Approach to Efficient LiDAR Multi-Task Perception

IROS 2024

Multi-task perception networks hold great potential as they can improve performance and computational efficiency compared to their single-task counterparts, facilitating online deployment. However, current multi-task architectures in point cloud perception combine multiple task-specific point cloud

Cited by 10SourceScholar
2024

AmodalSynthDrive: A Synthetic Amodal Perception Dataset for Autonomous Driving

RA-L 2024

Unlike humans, who can effortlessly estimate the entirety of objects even when partially occluded, modern computer vision algorithms still find this aspect extremely challenging. Leveraging this amodal perception for autonomous driving remains largely untapped due to the lack of suitable datasets. T

Cited by 16SourceScholar
2024

Automatic Target-Less Camera-LiDAR Calibration From Motion and Deep Point Correspondences

RA-L 2024

Sensor setups of robotic platforms commonly include both camera and LiDAR as they provide complementary information. However, fusing these two modalities typically requires a highly accurate calibration between them. In this letter, we propose MDPCalib which is a novel method for camera-LiDAR calibr

Cited by 16SourcecodeScholar
2024

BEVCar: Camera-Radar Fusion for BEV Map and Object Segmentation

IROS 2024poster

Semantic scene segmentation from a bird’s-eye-view (BEV) perspective plays a crucial role in facilitating planning and decision-making for mobile robots. Although recent vision-only methods have demonstrated notable advancements in performance, they often struggle under adverse illumination conditio…

Cited by 13SourcecodeScholar
2024

Bayesian Optimization for Sample-Efficient Policy Improvement in Robotic Manipulation

IROS 2024poster

Sample efficient learning of manipulation skills poses a major challenge in robotics. While recent approaches demonstrate impressive advances in the type of task that can be addressed and the sensing modalities that can be incorporated, they still require large amounts of training data. Especially w…

Cited by 1SourceScholar
2024

CenterGrasp: Object-Aware Implicit Representation Learning for Simultaneous Shape Reconstruction and 6-DoF Grasp Estimation

RA-L 2024

Reliable object grasping is a crucial capability for autonomous robots. However, many existing grasping approaches focus on general clutter removal without explicitly modeling objects and thus only relying on the visible local geometry. We introduce CenterGrasp, a novel framework that combines objec

Cited by 27SourceScholar
2024

Collaborative Dynamic 3D Scene Graphs for Automated Driving

ICRA 2024poster

Maps have played an indispensable role in enabling safe and automated driving. Although there have been many advances on different fronts ranging from SLAM to semantics, building an actionable hierarchical semantic representation of urban dynamic scenes and processing information from multiple agent…

Cited by 27SourcecodeScholar
2024

Compositional Servoing by Recombining Demonstrations

ICRA 2024poster

Learning-based manipulation policies from image inputs often show weak task transfer capabilities. In contrast, visual servoing methods allow efficient task transfer in high-precision scenarios while requiring only a few demonstrations. In this work, we present a framework that formulates the visual…

Cited by 0SourceScholar
2024

DITTO: Demonstration Imitation by Trajectory Transformation

IROS 2024poster

Teaching robots new skills quickly and conveniently is crucial for the broader adoption of robotic systems. In this work, we address the problem of one-shot imitation from a single human demonstration, given by an RGB-D video recording. We propose a two-stage process. In the first stage we extract t…

Cited by 16SourcecodeScholar
2024

Few-Shot Panoptic Segmentation With Foundation Models

ICRA 2024poster

Current state-of-the-art methods for panoptic segmentation require an immense amount of annotated training data that is both arduous and expensive to obtain posing a significant challenge for their widespread adoption. Concurrently, recent breakthroughs in visual representation learning have sparked…

Cited by 21SourcecodeScholar
2024

Hierarchical Open-Vocabulary 3D Scene Graphs for Language-Grounded Robot Navigation

RSS 2024poster

Recent open-vocabulary robot mapping methods enrich dense geometric maps with pre-trained visual-language features. While these maps allow for the prediction of point-wise saliency maps when queried for a certain language concept, large-scale environments and abstract queries beyond the object level…

2024

Language-Grounded Dynamic Scene Graphs for Interactive Object Search With Mobile Manipulation

RA-L 2024

To fully leverage the capabilities of mobile manipulation robots, it is imperative that they are able to autonomously execute long-horizon tasks in large unexplored environments. While large language models (LLMs) have shown emergent reasoning skills on arbitrary tasks, existing work primarily conce

Cited by 100SourcecodeScholar
2024

Learning Robotic Manipulation Policies from Point Clouds with Conditional Flow Matching

CoRL 2024poster

Learning from expert demonstrations is a popular approach to train robotic manipulation policies from limited data. However, imitation learning algorithms require a number of design choices ranging from the input modality, training objective, and 6-DoF end-effector pose representation. Diffusion-bas…

Cited by 14SourceScholar
2024

LetsMap: Unsupervised Representation Learning for Label-Efficient Semantic BEV Mapping

ECCV 2024poster

"Semantic Bird’s Eye View (BEV) maps offer a rich representation with strong occlusion reasoning for various decision making tasks in autonomous driving. However, most BEV mapping approaches employ a fully supervised learning paradigm that relies on large amounts of human-annotated BEV ground truth…

Cited by 1SourcePDFScholar
2024

Online Estimation of Articulated Objects with Factor Graphs using Vision and Proprioceptive Sensing

ICRA 2024poster

From dishwashers to cabinets, humans interact with articulated objects every day, and for a robot to assist in common manipulation tasks, it must learn a representation of articulation. Recent deep learning methods can provide powerful vision-based priors on the affordance of articulated objects fro…

Cited by 9SourcecodeScholar
2024

Progressive Multi-Modal Fusion for Robust 3D Object Detection

CoRL 2024poster

Multi-sensor fusion is crucial for accurate 3D object detection in autonomous driving, with cameras and LiDAR being the most commonly used sensors. However, existing methods perform sensor fusion in a single view by projecting features from both modalities either in Bird's Eye View (BEV) or Perspect…

Cited by 3SourceScholar
2024

Self-Supervised Representation Learning From Temporal Ordering of Automated Driving Sequences

RA-L 2024

Self-supervised feature learning enables perception systems to benefit from the vast raw data recorded by vehicle fleets worldwide. While video-level self-supervised approaches have shown strong generalizability on classification tasks, the potential to learn dense representations from sequential da

Cited by 14SourceScholar
2024

Syn-Mediverse: A Multimodal Synthetic Dataset for Intelligent Scene Understanding of Healthcare Facilities

RA-L 2024

Safety and efficiency are paramount in healthcare facilities where the lives of patients are at stake. Despite the adoption of robots to assist medical staff in challenging tasks such as complex surgeries, human expertise is still indispensable. The next generation of autonomous healthcare robots hi

Cited by 9SourceScholar
2024

The Art of Imitation: Learning Long-Horizon Manipulation Tasks From Few Demonstrations

RA-L 2024

Task Parametrized Gaussian Mixture Models (TP-GMM) are a sample-efficient method for learning object-centric robot manipulation tasks. However, there are several open challenges to applying TP-GMMs in the wild. In this work, we tackle three crucial challenges synergistically. First, end-effector vel

Cited by 15SourcecodeScholar
2023

CARTO: Category and Joint Agnostic Reconstruction of ARTiculated Objects

CVPR 2023poster

We present CARTO, a novel approach for reconstructing multiple articulated objects from a single stereo RGB observation. We use implicit object-centric representations and learn a single geometry and articulation decoder for multiple object categories. Despite training on multiple categories, our de…

2023

Catch Me if You Hear Me: Audio-Visual Navigation in Complex Unmapped Environments With Moving Sounds

RA-L 2023

Audio-visual navigation combines sight and hearing to navigate to a sound-emitting source in an unmapped environment. While recent approaches have demonstrated the benefits of audio input to detect and find the goal, they focus on clean and static sound sources and struggle to generalize to unheard

Cited by 54SourcecodeScholar
2023

CoDEPS: Online Continual Learning for Depth Estimation and Panoptic Segmentation

RSS 2023poster

Operating a robot in the open world requires a high level of robustness with respect to previously unseen environments. Optimally, the robot is able to adapt by itself to new conditions without human supervision, e.g., automatically adjusting its perception system to changing lighting conditions. In…

2023

EvCenterNet: Uncertainty Estimation for Object Detection Using Evidential Learning

IROS 2023poster

Uncertainty estimation is crucial in safety-critical settings such as automated driving as it provides valuable information for several downstream tasks including high-level decision making and path planning. In this work, we propose EvCenterNet, a novel uncertainty-aware 2D object detection framewo…

Cited by 7SourceScholar
2023

INoD: Injected Noise Discriminator for Self-Supervised Representation Learning in Agricultural Fields

RA-L 2023

Perception datasets for agriculture are limited both in quantity and diversity which hinders effective training of supervised learning approaches. Self-supervised learning techniques alleviate this problem, however, existing methods are not optimized for dense prediction tasks in agricultural domain

Cited by 11SourceScholar
2023

Learning Hierarchical Interactive Multi-Object Search for Mobile Manipulation

RA-L 2023

Existing object-search approaches enable robots to search through free pathways, however, robots operating in unstructured human-centered environments frequently also have to manipulate the environment to their needs. In this work, we introduce a novel interactive multi-object search task in which a

Cited by 33SourceScholar
2023

Learning and Aggregating Lane Graphs for Urban Automated Driving

CVPR 2023poster

Lane graph estimation is an essential and highly challenging task in automated driving and HD map learning. Existing methods using either onboard or aerial imagery struggle with complex lane topologies, out-of-distribution scenarios, or significant occlusions in the image space. Moreover, merging ov…

Cited by 29SourcePDFScholar
2023

PADLoC: LiDAR-Based Deep Loop Closure Detection and Registration Using Panoptic Attention

RA-L 2023

A key component of graph-based SLAM systems is the ability to detect loop closures in a trajectory to reduce the drift accumulated over time from the odometry. Most LiDAR-based methods achieve this goal by using only the geometric information, disregarding the semantics of the scene. In this work, w

Cited by 40SourcecodeScholar
2023

Self-Supervised Multi-Object Tracking for Autonomous Driving From Consistency Across Timescales

RA-L 2023

Self-supervised multi-object trackers have tremendous potential as they enable learning from raw domain-specific data. However, their re-identification accuracy still falls short compared to their supervised counterparts. We hypothesize that this drawback results from formulating self-supervised obj

Cited by 11SourceScholar
2023

SkyEye: Self-Supervised Bird's-Eye-View Semantic Mapping Using Monocular Frontal View Images

CVPR 2023poster

Bird's-Eye-View (BEV) semantic maps have become an essential component of automated driving pipelines due to the rich representation they provide for decision-making tasks. However, existing approaches for generating these maps still follow a fully supervised training paradigm and hence rely on larg…

Cited by 39SourcePDFScholar
2023

The Treachery of Images: Bayesian Scene Keypoints for Deep Policy Learning in Robotic Manipulation

RA-L 2023

In policy learning for robotic manipulation, sample efficiency is of paramount importance. Thus, learning and extracting more compact representations from camera observations is a promising avenue. However, current methods often assume full observability of the scene and struggle with scale invarian

Cited by 15SourcecodeScholar
2022

Amodal Panoptic Segmentation

CVPR 2022poster

Humans have the remarkable ability to perceive objects as a whole, even when parts of them are occluded. This ability of amodal perception forms the basis of our perceptual and cognitive understanding of our world. To enable robots to reason with this capability, we formulate and propose a novel tas…

Cited by 48PDFScholar
2022

Correct Me If I am Wrong: Interactive Learning for Robotic Manipulation

RA-L 2022

Learning to solve complex manipulation tasks from visual observations is a dominant challenge for real-world robot learning. Although deep reinforcement learning algorithms have recently demonstrated impressive results in this context, they still require an impractical amount of time-consuming trial

Cited by 48SourceScholar
2022

Kineverse: A Symbolic Articulation Model Framework for Model-Agnostic Mobile Manipulation

RA-L 2022

Service robots in the future need to execute abstract instructions such as “fetch the milk from the fridge”. To translate such instructions into actionable plans, robots require in-depth background knowledge. With regards to interactions with doors and drawers, robots require articulation models tha

Cited by 17SourceScholar
2022

Panoptic Nuscenes: A Large-Scale Benchmark for LiDAR Panoptic Segmentation and Tracking

RA-L 2022

Panoptic scene understanding and tracking of dynamic agents are essential for robots and automated vehicles to navigate in urban environments. As LiDARs provide accurate illumination-independent geometric depictions of the scene, performing these tasks using LiDAR point clouds provides reliable pred

Cited by 243SourceScholar
2022

Robot Skill Adaptation via Soft Actor-Critic Gaussian Mixture Models

ICRA 2022poster

AA core challenge for an autonomous agent acting in the real world is to adapt its repertoire of skills to cope with its noisy perception and dynamics. To scale learning of skills to long-horizon tasks, robots should be able to learn and later refine their skills in a structured manner through traje…

Cited by 23SourceScholar
2021

Learning Kinematic Feasibility for Mobile Manipulation Through Deep Reinforcement Learning

RA-L 2021

Mobile manipulation tasks remain one of the critical challenges for the widespread adoption of autonomous robots in both service and industrial scenarios. While planning approaches are good at generating feasible whole-body robot trajectories, they struggle with dynamic environments as well as the i

Cited by 59SourcecodeScholar
2021

There Is More Than Meets the Eye: Self-Supervised Multi-Object Detection and Tracking With Sound by Distilling Multimodal Knowledge

CVPR 2021poster

Attributes of sound inherent to objects can provide valuable cues to learn rich representations for object detection and tracking. Furthermore, the co-occurrence of audiovisual events in videos can be exploited to localize objects over the image field by solely monitoring the sound in the environmen…

Cited by 97PDFScholar
2019

Robot Localization in Floor Plans Using a Room Layout Edge Extraction Network

IROS 2019poster

Indoor localization is one of the crucial enablers for deployment of service robots. Although several successful techniques for indoor localization have been proposed, the majority of them relies on maps generated from data gathered with the same sensor modality used for localization. Typically, ted…

Cited by 56SourceScholar
2017

AdapNet: Adaptive semantic segmentation in adverse environmental conditions

ICRA 2017poster

Robust scene understanding of outdoor environments using passive optical sensors is a onerous and essential task for autonomous navigation. The problem is heavily characterized by changing environmental conditions throughout the day and across seasons. Robots should be equipped with models that are…

Cited by 265SourceScholar
2017

SMSnet: Semantic motion segmentation using deep convolutional neural networks

IROS 2017poster

Interpreting the semantics and motion of objects are prerequisites for autonomous robots that enable them to reason and operate in dynamic real-world environments. Existing approaches that tackle the problem of semantic motion segmentation consist of long multistage pipelines and typically require s…

Cited by 91SourceScholar
2016

Autonomous indoor robot navigation using a sketch interface for drawing maps and routes

ICRA 2016

Hand-Drawn sketches are natural means by which abstract descriptions of environments can be provided. They represent weak prior information about the scene, thereby enabling a robot to perform autonomous navigation and exploration when a full metrical description of the environment is not available

Cited by 53SourceScholar
2016

Deep learning for human part discovery in images

ICRA 2016

This paper addresses the problem of human body part segmentation in conventional RGB images, which has several applications in robotics, such as learning from demonstration and human-robot handovers. The proposed solution is based on Convolutional Neural Networks (CNNs). We present a network archite

Cited by 107SourceScholar