← Search

Tushar Nagarajan

31 accepted papers

2026

Cross-Embodiment Robot Foundation World Models with Latent Actions

ICML 2026poster

The diversity of robot embodiments and action spaces makes it challenging to build robot world models that generalize across different embodiments. We introduce a Latent Action Conditioned Robot World Model (LAC-WM), which operates within a learned unified latent action space shared across diverse e…

Cited by 0SourceScholar
2026

Learning Latent Action World Models In The Wild

ICML 2026poster

Agents that can reason and plan in the real world must be able to predict the consequences of their actions. World models possess this capability but require action annotations that can be complex to obtain at scale. Latent action models address this issue by learning an action space from videos alo…

Cited by 0SourceScholar
2025

BIMBA: Selective-Scan Compression for Long-Range Video Question Answering

CVPR 2025poster

Video Question Answering (VQA) in long videos poses the key challenge of extracting relevant information and modeling long-range dependencies from many redundant frames. The self-attention mechanism provides a general solution for sequence modeling, but it has a prohibitive cost when applied to a ma…

2025

EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception

ICCV 2025poster

Modern perception models, particularly those designed for multisensory egocentric tasks, have achieved remarkable performance but often come with substantial computational costs. These high demands pose challenges for real-world deployment, especially in resource-constrained environments. In this pa…

Cited by 0SourcePDFScholar
2025

ExpertAF: Expert Actionable Feedback from Video

CVPR 2025poster

Feedback is essential for learning a new skill or improving one's current skill-level. However, current methods for skill-assessment from video only provide scores or compare demonstrations, leaving the burden of knowing what to do differently on the user. We introduce a novel method to generate act…

Cited by 3SourcePDFScholar
2025

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

NeurIPS 2025spotlight

Vision-language models are integral to computer vision research, yet many high-performing models remain closed-source, obscuring their data, design and training recipe. The research community has responded by using distillation from black-box models to label training data, achieving strong benchmark…

Cited by 0SourcecodeScholar
2025

Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos

ICCV 2025poster

We introduce Switch-a-View, a model that learns to automatically select the viewpoint to display at each timepoint when creating a how-to video. The key insight of our approach is how to train such a model from unlabeled--but human-edited--video samples. We pose a pretext task that pseudo-labels seg…

Cited by 0SourcePDFScholar
2025

VEDIT: Latent Prediction Architecture For Procedural Video Representation Learning

ICLR 2025poster

Procedural video representation learning is an active research area where the objective is to learn an agent which can anticipate and forecast the future given the present video input, typically in conjunction with textual annotations. Prior works often rely on large-scale pretraining of visual enco…

Cited by 1SourcePDFScholar
2025

VITED: Video Temporal Evidence Distillation

CVPR 2025poster

We investigate complex video question answering via chain-of-evidence reasoning --- identifying sequences of temporal spans from multiple relevant parts of the video, together with visual evidence within them.Existing models struggle with multi-step reasoning as they uniformly sample a fixed number…

Cited by 1SourcePDFScholar
2025

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos

CVPR 2025highlight

Given a multi-view video, which viewpoint is most informative for a human observer? Existing methods rely on heuristics or expensive "best-view" supervision to answer this question, limiting their applicability. We propose a weakly supervised approach that leverages language accompanying an instruct…

Cited by 0SourcePDFScholar
2024

"Propose, Assess, Search: Harnessing LLMs for Goal-Oriented Planning in Instructional Videos"

ECCV 2024oral

"Goal-oriented planning, or anticipating a series of actions that transition an agent from its current state to a predefined objective, is crucial for developing intelligent assistants aiding users in daily procedural tasks. The problem presents significant challenges due to the need for comprehensi…

Cited by 2SourcePDFScholar
2024

AMEGO: Active Memory from long EGOcentric videos

ECCV 2024poster

"Egocentric videos provide a unique perspective into individuals’ daily experiences, yet their unstructured nature presents challenges for perception. In this paper, we introduce , a novel approach aimed at enhancing the comprehension of very-long egocentric videos. Inspired by the human’s ability t…

Cited by 8SourcePDFScholar
2024

AnyMAL: An Efficient and Scalable Any-Modality Augmented Language Model

EMNLP 2024industry

We present Any-Modality Augmented Language Model (AnyMAL), a unified model that reasons over diverse input modality signals (i.e. text, image, video, audio, IMU motion sensor), and generates textual responses. AnyMAL inherits the powerful text-based reasoning abilities of the state-of-the-art LLMs i…

2024

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

CVPR 2024poster

We present Ego-Exo4D a diverse large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g. sports music dance bike repair). 740 participants from 13 cities worldwide perform…

2024

Video ReCap: Recursive Captioning of Hour-Long Videos

CVPR 2024poster

Most video captioning models are designed to process short video clips of few seconds and output text describing low-level visual concepts (e.g. objects scenes atomic actions). However most real-world videos last for minutes or hours and have a complex hierarchical structure spanning different tempo…

2023

Ego4D Goal-Step: Toward Hierarchical Understanding of Procedural Activities

NeurIPS 2023spotlight

Human activities are goal-oriented and hierarchical, comprising primary goals at the top level, sequences of steps and substeps in the middle, and atomic actions at the lowest level. Recognizing human activities thus requires relating atomic actions and steps to their functional objectives (what the…

Cited by 35SourcePDFScholar
2023

EgoDistill: Egocentric Head Motion Distillation for Efficient Video Understanding

NeurIPS 2023poster

Recent advances in egocentric video understanding models are promising, but their heavy computational expense is a barrier for many real-world applications. To address this challenge, we propose EgoDistill, a distillation-based approach that learns to reconstruct heavy ego-centric video clip feature…

Cited by 25SourcePDFScholar
2023

EgoEnv: Human-centric environment representations from egocentric video

NeurIPS 2023oral

First-person video highlights a camera-wearer's activities in the context of their persistent environment. However, current video understanding approaches reason over visual features from short video clips that are detached from the underlying physical space and capture only what is immediately vis…

Cited by 20SourcePDFScholar
2023

HT-Step: Aligning Instructional Articles with How-To Videos

NeurIPS 2023poster

We introduce HT-Step, a large-scale dataset containing temporal annotations of instructional article steps in cooking videos. It includes 122k segment-level annotations over 20k narrated videos (approximately 2.3k hours) of the HowTo100M dataset. Each annotation provides a temporal interval, and a c…

Cited by 18SourcePDFScholar
2022

Ego4D: Around the World in 3,000 Hours of Egocentric Video

CVPR 2022oral

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countri…

Cited by 1162PDFcodeScholar
2022

Environment Predictive Coding for Visual Navigation

ICLR 2022poster

We introduce environment predictive coding, a self-supervised approach to learn environment-level representations for embodied agents. In contrast to prior work on self-supervised learning for individual images, we aim to encode a 3D environment using a series of images observed by an agent moving i…

Cited by 8SourcePDFScholar
2021

Differentiable Causal Discovery Under Unmeasured Confounding

AISTATS 2021poster

The data drawn from biological, economic, and social systems are often confounded due to the presence of unmeasured variables. Prior work in causal discovery has focused on discrete search procedures for selecting acyclic directed mixed graphs (ADMGs), specifically ancestral ADMGs, that encode ordin…

2021

Ego-Exo: Transferring Visual Representations From Third-Person to First-Person Videos

CVPR 2021poster

We introduce an approach for pre-training egocentric video models using large-scale third-person video datasets. Learning from purely egocentric data is limited by low dataset scale and diversity, while using purely exocentric (third-person) data introduces a large domain mismatch. Our idea is to di…

Cited by 98PDFcodeScholar
2021

Shaping embodied agent behavior with activity-context priors from egocentric video

NeurIPS 2021spotlight

Complex physical tasks entail a sequence of object interactions, each with its own preconditions -- which can be difficult for robotic agents to learn efficiently solely through their own experience. We introduce an approach to discover activity-context priors from in-the-wild egocentric video captu…

Cited by 18SourcePDFScholar
2020

Ego-Topo: Environment Affordances From Egocentric Video

CVPR 2020oral

First-person video naturally brings the use of a physical environment to the forefront, since it shows the camera wearer interacting fluidly in a space based on his intentions. However, current methods largely separate the observed actions from the persistent space itself. We introduce a model for e…

Cited by 152PDFcodeScholar
2020

Learning Affordance Landscapes for Interaction Exploration in 3D Environments

NeurIPS 2020spotlight

Embodied agents operating in human spaces must be able to master how their environment works: what objects can the agent use, and how can it use them? We introduce a reinforcement learning approach for exploration for interaction, whereby an embodied agent autonomously discovers the affordance lands…

Cited by 92SourcePDFScholar
2018

Attributes as Operators: Factorizing Unseen Attribute-Object Compositions

ECCV 2018poster

We present a new approach to modeling visual attributes. Prior work casts attributes in a similar role as objects, learning a latent representation where properties (e.g., sliced) are recognized by classifiers much in the way objects (e.g., apple) are. However, this common approach fails to separate…

2018

BlockDrop: Dynamic Inference Paths in Residual Networks

CVPR 2018poster

Very deep convolutional neural networks offer excellent recognition results, yet their computational expense limits their impact for many real-world applications. We introduce BlockDrop, an approach that learns to dynamically choose which layers of a deep network to execute during inference so as t…