← Search

Sanjay Haresh

9 accepted papers

2026

Notes-To-Self: Scratchpad Augmented VLAs for Memory Dependent Manipulation Tasks

ICRA 2026poster

Many dexterous manipulation tasks are non-markovian in nature, yet little attention has been paid to this fact in the recent upsurge of the vision-language-action (VLA) paradigm. Although they are successful in bringing internet-scale semantic understanding to robotics, existing VLAs are primarily "…

2025

Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?

NeurIPS 2025poster

Multi-modal Large Language Models (LLM) have advanced conversational abilities but struggle with providing live, interactive step-by-step guidance, a key capability for future AI assistants. Effective guidance requires not only delivering instructions but also detecting their successful execution, a…

Cited by 0SourcecodeScholar
2025

Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models

CoRL 2025poster

Vision-Language-Action (VLA) models offer a pivotal approach to learning robotic manipulation at scale by repurposing large pre-trained Vision-Language-Models (VLM) to output robotic actions. However, adapting VLMs for robotic domains comes with an unnecessarily high computational cost, which we att…

Cited by 0SourceScholar
2024

ClevrSkills: Compositional Language And Visual Reasoning in Robotics

NeurIPS 2024poster

Robotics tasks are highly compositional by nature. For example, to perform a high-level task like cleaning the table a robot must employ low-level capabilities of moving the effectors to the objects on the table, pick them up and then move them off the table one-by-one, while re-evaluating the conse…

2024

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

CVPR 2024poster

We present Ego-Exo4D a diverse large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g. sports music dance bike repair). 740 participants from 13 cities worldwide perform…

2024

Habitat Synthetic Scenes Dataset (HSSD-200): An Analysis of 3D Scene Scale and Realism Tradeoffs for ObjectGoal Navigation

CVPR 2024poster

We contribute the Habitat Synthetic Scene Dataset a dataset of 211 high-quality 3D scenes and use it to test navigation agent generalization to realistic 3D environments. Our dataset represents real interiors and contains a diverse set of 18656 models of real-world objects. We investigate the impact…

Cited by 53SourcePDFScholar
2022

Timestamp-Supervised Action Segmentation with Graph Convolutional Networks

IROS 2022poster

We introduce a novel approach for temporal activity segmentation with timestamp supervision. Our main contribution is a graph convolutional network, which is learned in an end-to-end manner to exploit both frame features and connections between neighboring frames to generate dense framewise labels f…

Cited by 18SourceScholar
2022

Unsupervised Action Segmentation by Joint Representation Learning and Online Clustering

CVPR 2022poster

We present a novel approach for unsupervised activity segmentation which uses video frame clustering as a pretext task and simultaneously performs representation learning and online clustering. This is in contrast with prior works where representation learning and clustering are often performed sequ…

Cited by 70PDFcodeScholar
2021

Learning by Aligning Videos in Time

CVPR 2021poster

We present a self-supervised approach for learning video representations using temporal video alignment as a pretext task, while exploiting both frame-level and video-level information. We leverage a novel combination of temporal alignment loss and temporal regularization terms, which can be used as…

Cited by 86PDFScholar