← Search

Nakul Agarwal

13 accepted papers

2026

MERGE: Guided Vision-Language Models for Multi-Actor Event Reasoning and Grounding in Human–Robot Interaction

ICRA 2026poster

We introduce MERGE, a system for situational grounding of actors, objects, and events in dynamic human–robot group interactions. Effective collaboration in such settings requires consistent situational awareness, built on persistent representations of people and objects and an episodic abstraction o…

2025

Overcoming Multi-step Complexity in Multimodal Theory-of-Mind Reasoning: A Scalable Bayesian Planner

ICML 2025spotlight

Theory-of-mind (ToM) enables humans to infer mental states—such as beliefs, desires, and intentions—forming the foundation of social cognition. Existing computational ToM methods rely on structured workflows with ToM-specific priors or deep model fine-tuning but struggle with scalability in multimod…

Cited by 0SourcePDFScholar
2024

AntGPT: Can Large Language Models Help Long-term Action Anticipation from Videos?

ICLR 2024poster

Can we better anticipate an actor’s future actions (e.g. mix eggs) by knowing what commonly happens after the current action (e.g. crack eggs)? What if the actor also shares the goal (e.g. make fried rice) with us? The long-term action anticipation (LTA) task aims to predict an actor’s future behavi…

2024

Can't Make an Omelette Without Breaking Some Eggs: Plausible Action Anticipation Using Large Video-Language Models

CVPR 2024poster

We introduce PlausiVL a large video-language model for anticipating action sequences that are plausible in the real-world. While significant efforts have been made towards anticipating future actions prior approaches do not take into account the aspect of plausibility in an action sequence. To addre…

Cited by 18SourcePDFScholar
2024

Disentangled Neural Relational Inference for Interpretable Motion Prediction

RA-L 2024

Effective interaction modeling and behavior prediction of dynamic agents play a significant role in interactive motion planning for autonomous robots. Although existing methods have improved prediction accuracy, few research efforts have been devoted to enhancing prediction model interpretability an

Cited by 9SourceScholar
2024

M2D2M: Multi-Motion Generation from Text with Discrete Diffusion Models

ECCV 2024poster

"We introduce the Multi-Motion Discrete Diffusion Models (M2D2M), a novel approach for human motion generation from textual descriptions of multiple actions, utilizing the strengths of discrete diffusion models. This approach adeptly addresses the challenge of generating multi-motion sequences, ensu…

Cited by 11SourcePDFScholar
2024

Uncertainty-aware Action Decoupling Transformer for Action Anticipation

CVPR 2024highlight

Human action anticipation aims at predicting what people will do in the future based on past observations. In this paper we introduce Uncertainty-aware Action Decoupling Transformer (UADT) for action anticipation. Unlike existing methods that directly predict action in a verb-noun pair format we dec…

Cited by 11SourcePDFScholar
2024

Vamos: Versatile Action Models for Video Understanding

ECCV 2024poster

"What makes good representations for video understanding, such as anticipating future activities, or answering video-conditioned questions? While earlier approaches focus on end-to-end learning directly from video pixels, we propose to revisit text-based representations, such as general-purpose vide…

2023

AdamsFormer for Spatial Action Localization in the Future

CVPR 2023poster

Predicting future action locations is vital for applications like human-robot collaboration. While some computer vision tasks have made progress in predicting human actions, accurately localizing these actions in future frames remains an area with room for improvement. We introduce a new task called…

2023

Latency Matters: Real-Time Action Forecasting Transformer

CVPR 2023highlight

We present RAFTformer, a real-time action forecasting transformer for latency aware real-world action forecasting applications. RAFTformer is a two-stage fully transformer based architecture which consists of a video transformer backbone that operates on high resolution, short range clips and a head…

2023

Weakly-Supervised Action Segmentation and Unseen Error Detection in Anomalous Instructional Videos

ICCV 2023poster

We present a novel method for weakly-supervised action segmentation and unseen error detection in anomalous instructional videos. In the absence of an appropriate dataset for this task, we introduce the Anomalous Toy Assembly (ATA) dataset, which comprises 1152 untrimmed videos of 32 participants as…

Cited by 19PDFScholar
2022

Weakly-Supervised Online Action Segmentation in Multi-View Instructional Videos

CVPR 2022poster

This paper addresses a new problem of weakly-supervised online action segmentation in instructional videos. We present a framework to segment streaming videos online at test time using Dynamic Programming and show its advantages over greedy sliding window approach. We improve our framework by introd…

Cited by 26PDFScholar