← Search

Apoorv Khandelwal

4 accepted papers

2022

A-OKVQA: A Benchmark for Visual Question Answering Using World Knowledge

ECCV 2022poster

"The Visual Question Answering (VQA) task aspires to provide a meaningful testbed for the development of AI models that can jointly reason over visual and natural language inputs. Despite a proliferation of VQA datasets, this goal is hindered by a set of common limitations. These include a reliance…

2022

Simple but Effective: CLIP Embeddings for Embodied AI

CVPR 2022poster

Contrastive language image pretraining (CLIP) encoders have been shown to be beneficial for a range of visual tasks from classification and detection to captioning and image manipulation. We investigate the effectiveness of CLIP visual backbones for Embodied AI tasks. We build incredibly simple base…

Cited by 252PDFcodeScholar
2021

Who's Waldo? Linking People Across Text and Images

ICCV 2021poster

We present a task and benchmark dataset for person-centric visual grounding, the problem of linking between people named in a caption and people pictured in an image. In contrast to prior work in visual grounding, which is predominantly object-based, our new task masks out the names of people in cap…

Cited by 22PDFcodeScholar