← Search

Anil Batra

4 accepted papers

2025

CAST: Cross-modal Alignment Similarity Test for Vision Language Models

COLING 2025main

Vision Language Models (VLMs) are typically evaluated with Visual Question Answering (VQA) tasks which assess a model’s understanding of scenes. Good VQA performance is taken as evidence that the model will perform well on a broader range of tasks that require both visual and language inputs. Howeve…

2025

Predicting Implicit Arguments in Procedural Video Instructions

ACL 2025long

Procedural texts help AI enhance reasoning about context and action sequences. Transforming these into Semantic Role Labeling (SRL) improves understanding of individual steps by identifying predicate-argument structure like verb,what,where/with. Procedural instructions are highly elliptic, for insta…

2023

Image generation with shortest path diffusion

ICML 2023poster

The field of image generation has made significant progress thanks to the introduction of Diffusion Models, which learn to progressively reverse a given image corruption. Recently, a few studies introduced alternative ways of corrupting images in Diffusion Models, with an emphasis on blurring. Howev…

2019

Improved Road Connectivity by Joint Learning of Orientation and Segmentation

CVPR 2019poster

Road network extraction from satellite images often produce fragmented road segments leading to road maps unfit for real applications. Pixel-wise classification fails to predict topologically correct and connected road masks due to the absence of connectivity supervision and difficulty in enforcing…

Cited by 237PDFScholar