← Search

Makarand Tapaswi

26 accepted papers

2025

IdentifyMe: A Challenging Long-Context Mention Resolution Benchmark for LLMs

NAACL 2025short

Recent evaluations of LLMs on coreference resolution have revealed that traditional output formats and evaluation metrics do not fully capture the models’ referential understanding. To address this, we introduce IdentifyMe, a new benchmark for mention resolution presented in a multiple-choice questi…

2025

The Sound of Water: Inferring Physical Properties from Pouring Liquids

ICASSP 2025accepted

We study the connection between audio-visual observations and the underlying physics of a mundane yet intriguing everyday activity: pouring liquids. Given only the sound of liquid pouring into a container, our objective is to automatically infer physical properties such as the liquid level, the shap…

Cited by 0SourceScholar
2025

VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailment

CVPR 2025poster

A fundamental aspect of compositional reasoning in a video is associating people and their actions across time. Recent years have seen great progress in general-purpose vision/video models and a move towards long-video understanding. While exciting, we take a step back and ask: are today's models go…

2024

MICap: A Unified Model for Identity-Aware Movie Descriptions

CVPR 2024poster

Characters are an important aspect of any storyline and identifying and including them in descriptions is necessary for story understanding. While previous work has largely ignored identity and generated captions with someone (anonymized names) recent work formulates id-aware captioning as a fill-in…

Cited by 3SourcePDFScholar
2024

Major Entity Identification: A Generalizable Alternative to Coreference Resolution

EMNLP 2024main

The limited generalization of coreference resolution (CR) models has been a major bottleneck in the task’s broad application. Prior work has identified annotation differences, especially for mention detection, as one of the main reasons for the generalization gap and proposed using additional annota…

2023

How You Feelin'? Learning Emotions and Mental States in Movie Scenes

CVPR 2023poster

Movie story analysis requires understanding characters' emotions and mental states. Towards this goal, we formulate emotion understanding as predicting a diverse and multi-label set of emotions at the level of a movie scene and for each character. We propose EmoTx, a multimodal Transformer-based arc…

2023

Test of Time: Instilling Video-Language Models With a Sense of Time

CVPR 2023poster

Modelling and understanding time remains a challenge in contemporary video understanding models. With language emerging as a key driver towards powerful generalization, it is imperative for foundational video-language models to have a sense of time. In this paper, we consider a specific aspect of te…

2022

Instruction-driven history-aware policies for robotic manipulations

CoRL 2022oral

In human environments, robots are expected to accomplish a variety of manipulation tasks given simple natural language instructions. Yet, robotic manipulation is extremely challenging as it requires fine-grained motor control, long-term memory as well as generalization to previously unseen tasks and…

Cited by 117SourcecodeScholar
2022

Language Conditioned Spatial Relation Reasoning for 3D Object Grounding

NeurIPS 2022accept

Localizing objects in 3D scenes based on natural language requires understanding and reasoning about spatial relations. In particular, it is often crucial to distinguish similar objects referred by the text, such as "the left most chair" and "a chair next to the window". In this work we propose a la…

2022

Learning from Unlabeled 3D Environments for Vision-and-Language Navigation

ECCV 2022poster

"In vision-and-language navigation (VLN), an embodied agent is required to navigate in realistic 3D environments following natural language instructions. One major bottleneck for existing VLN approaches is the lack of sufficient training data, resulting in unsatisfactory generalization to unseen env…

2022

Think Global, Act Local: Dual-Scale Graph Transformer for Vision-and-Language Navigation

CVPR 2022oral

Following language instructions to navigate in unseen environments is a challenging problem for autonomous embodied agents. The agent not only needs to ground languages in visual scenes, but also should explore the environment to reach its target. In this work, we propose a dual-scale graph transfor…

Cited by 181PDFScholar
2021

Airbert: In-Domain Pretraining for Vision-and-Language Navigation

ICCV 2021poster

Vision-and-language navigation (VLN) aims to enable embodied agents to navigate in realistic environments using natural language instructions. Given the scarcity of domain-specific training data and the high diversity of image and language inputs, the generalization of VLN agents to unseen environme…

Cited by 169PDFcodeScholar
2020

Learning Object Manipulation Skills via Approximate State Estimation from Real Videos

CoRL 2020

Humans are adept at learning new tasks by watching a few instructional videos. On the other hand, robots that learn new actions either require a lot of effort through trial and error, or use expert demonstrations that are challenging to obtain. In this paper, we explore a method that facilitates lea

Cited by 0SourcePDFScholar
2019

HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

ICCV 2019poster

Learning text-video embeddings usually requires a dataset of video clips with manually provided captions. However, such datasets are expensive and time consuming to create and therefore difficult to obtain on a large scale. In this work, we propose instead to learn such embeddings from video data wi…

Cited by 1412PDFScholar
2018

MovieGraphs: Towards Understanding Human-Centric Situations From Videos

CVPR 2018poster

There is growing interest in artificial intelligence to build socially intelligent robots. This requires machines to have the ability to "read" people's emotions, motivations, and other factors that affect behavior. Towards this goal, we introduce a novel dataset called MovieGraphs which provides de…

Cited by 179SourcePDFScholar
2017

Situation Recognition With Graph Neural Networks

ICCV 2017poster

We address the problem of recognizing situations in images. Given an image, the task is to predict the most salient verb (action), and fill its semantic roles such as who is performing the action, what is the source and target of the action, etc. Different verbs have different roles (e.g. attacking…

Cited by 142PDFScholar
2016

MovieQA: Understanding Stories in Movies Through Question-Answering

CVPR 2016spotlight

We introduce the MovieQA dataset which aims to evaluate automatic story comprehension from both video and text. The dataset consists of 14,944 questions about 408 movies with high semantic diversity. The questions range from simpler "Who" did "What" to "Whom", to "Why" and "How" certain events occur…

Cited by 875PDFScholar
2016

Recovering the Missing Link: Predicting Class-Attribute Associations for Unsupervised Zero-Shot Learning

CVPR 2016accepted

Collecting training images for all visual categories is not only expensive but also impractical. Zero-shot learning (ZSL), especially using attributes, offers a pragmatic solution to this problem. However, at test time most attribute-based methods require a full description of attribute associations…

Cited by 126SourcePDFScholar