← Search

Ashutosh Saxena

12 accepted papers

2017

Deep multimodal embedding: Manipulating novel objects with point-clouds, language and trajectories

ICRA 2017poster

A robot operating in a real-world environment needs to perform reasoning over a variety of sensor modalities such as vision, language and motion trajectories. However, it is extremely challenging to manually design features relating such disparate modalities. In this work, we introduce an algorithm…

Cited by 39SourceScholar
2017

Learning to represent haptic feedback for partially-observable tasks

ICRA 2017poster

The sense of touch, being the earliest sensory system to develop in a human body [1], plays a critical part of our daily interaction with the environment. In order to successfully complete a task, many manipulation interactions require incorporating haptic feedback. However, manually designing a fee…

Cited by 34SourceScholar
2016

Learning Transferrable Representations for Unsupervised Domain Adaptation

NeurIPS 2016poster

Supervised learning with large scale labelled datasets and deep layered models has caused a paradigm shift in diverse areas in learning and recognition. However, this approach still suffers from generalization issues under the presence of a domain shift between the training and the test data distrib…

Cited by 345SourcePDFScholar
2016

Recurrent Neural Networks for driver activity anticipation via sensory-fusion architecture

ICRA 2016

Anticipating the future actions of a human is a widely studied problem in robotics that requires spatio-temporal reasoning. In this work we propose a deep learning approach for anticipation in sensory-rich robotics applications. We introduce a sensory-fusion architecture which jointly learns to anti

Cited by 274SourceScholar
2016

Structural-RNN: Deep Learning on Spatio-Temporal Graphs

CVPR 2016oral

Deep Recurrent Neural Network architectures, though remarkably capable at modeling sequences, lack an intuitive high-level spatio-temporal structure. That is while many problems in computer vision inherently have an underlying high-level structure and can benefit from it. Spatio-temporal graphs are…

Cited by 1477PDFcodeScholar
2016

Watch-Bot: Unsupervised learning for reminding humans of forgotten actions

ICRA 2016

We present a robotic system that watches a human using a Kinect v2 RGB-D sensor, detects what he forgot to do while performing an activity, and if necessary reminds the person using a laser pointer to point out the related object. Our simple setup can be easily deployed on any assistive robot. Our a

Cited by 17SourceScholar
2015

Car That Knows Before You Do: Anticipating Maneuvers via Learning Temporal Driving Models

ICCV 2015poster

Advanced Driver Assistance Systems (ADAS) have made driving safer over the last decade. They prepare vehicles for unsafe road conditions and alert drivers if they perform a dangerous maneuver. However, many accidents are unavoidable because by the time drivers are alerted, it is already too late.…

Cited by 350PDFScholar
2015

PlanIt: A crowdsourcing approach for learning to plan paths from large scale preference feedback

ICRA 2015poster

We consider the problem of learning user preferences over robot trajectories for environments rich in objects and humans. This is challenging because the criterion defining a good trajectory varies with users, tasks and interactions in the environment. We represent trajectory preferences using a cos…

Cited by 33SourceScholar
2015

Watch-n-Patch: Unsupervised Understanding of Actions and Relations

CVPR 2015poster

We focus on modeling human activities comprising multiple actions in a completely unsupervised setting. Our model learns the high-level action co-occurrence and temporal relations between the actions in the activity video. We consider the video as a sequence of short-term action clips, called action…

Cited by 182SourcePDFScholar