← Search

Ashish Tawari

5 accepted papers

2025

CoLLM: A Large Language Model for Composed Image Retrieval

CVPR 2025poster

Composed Image Retrieval (CIR) is a complex task that aims to retrieve images based on a multimodal query. Typical training data consists of triplets containing a reference image, a textual description of desired modifications, and the target image, which are expensive and time-consuming to acquire.…

2024

Open Vocabulary Multi-Label Video Classification

ECCV 2024poster

"Pre-trained vision-language models (VLMs) have enabled significant progress in open vocabulary computer vision tasks such as image classification, object detection and image segmentation. Some recent works have focused on extending VLMs to open vocabulary single label action classification in video…

Cited by 2SourcePDFScholar
2020

Interaction Graphs for Object Importance Estimation in On-road Driving Videos

ICRA 2020poster

A vehicle driving along the road is surrounded by many objects, but only a small subset of them influence the driver's decisions and actions. Learning to estimate the importance of each object on the driver's real-time decision-making may help better understand human driving behavior and lead to mor…

Cited by 34SourceScholar
2019

Grounding Human-To-Vehicle Advice for Self-Driving Vehicles

CVPR 2019poster

Recent success suggests that deep neural control networks are likely to be a key component of self-driving vehicles. These networks are trained on large datasets to imitate human actions, but they lack semantic understanding of image contents. This makes them brittle and potentially unsafe in situat…

Cited by 131PDFScholar