← Search

Sateesh Kumar

6 accepted papers

2026

MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos

ICRA 2026poster

We aim to enable humanoid robots to efficiently solve new manipulation tasks from a few video examples. In-context learning (ICL) is a promising framework for achieving this goal due to its test-time data efficiency and rapid adaptability. However, current ICL methods rely on labor-intensive teleope…

2026

Searching in Space and Time: Unified Memory-Action Loops for Open-World Object Retrieval

ICRA 2026poster

Service robots must retrieve objects in dynamic, open-world settings where requests may reference attributes (“the red mug”), spatial context (“the mug on the table”), or past states (“the mug that was here yesterday”). Existing approaches capture only parts of this problem: scene graphs capture spa…

2025

COLLAGE: Adaptive Fusion-based Retrieval for Augmented Policy Learning

CoRL 2025poster

In this work, we study the problem of data retrieval for few-shot imitation learning: select data from a large dataset to train a performant policy for a specific task, given only a few target demonstrations. Prior methods retrieve data using a single-feature distance heuristic, assuming that the be…

Cited by 0SourcecodeScholar
2022

Graph Inverse Reinforcement Learning from Diverse Videos

CoRL 2022oral

Research on Inverse Reinforcement Learning (IRL) from third-person videos has shown encouraging results on removing the need for manual reward design for robotic tasks. However, most prior works are still limited by training from a relatively restricted domain of videos. In this paper, we argue that…

Cited by 57SourcecodeScholar
2022

Unsupervised Action Segmentation by Joint Representation Learning and Online Clustering

CVPR 2022poster

We present a novel approach for unsupervised activity segmentation which uses video frame clustering as a pretext task and simultaneously performs representation learning and online clustering. This is in contrast with prior works where representation learning and clustering are often performed sequ…

Cited by 70PDFcodeScholar
2021

Learning by Aligning Videos in Time

CVPR 2021poster

We present a self-supervised approach for learning video representations using temporal video alignment as a pretext task, while exploiting both frame-level and video-level information. We leverage a novel combination of temporal alignment loss and temporal regularization terms, which can be used as…

Cited by 86PDFScholar