← Search

Karan Desai

7 accepted papers

2023

Hyperbolic Image-text Representations

ICML 2023poster

Visual and linguistic concepts naturally organize themselves in a hierarchy, where a textual concept "dog" entails all images that contain dogs. Despite being intuitive, current large-scale vision and language models such as CLIP do not explicitly capture such hierarchy. We propose MERU, a contrasti…

2023

Learning Visual Representations via Language-Guided Sampling

CVPR 2023poster

Although an object may appear in numerous contexts, we often describe it in a limited number of ways. Language allows us to abstract away visual variation to represent and communicate concepts. Building on this intuition, we propose an alternative approach to visual representation learning: using la…

2021

CASTing Your Model: Learning To Localize Improves Self-Supervised Representations

CVPR 2021poster

Recent advances in self-supervised learning (SSL) have largely closed the gap with supervised ImageNet pretraining. Despite their success these methods have been primarily applied to unlabeled ImageNet images, and show marginal gains when trained on larger sets of uncurated images. We hypothesize th…

Cited by 101PDFcodeScholar
2021

RedCaps: Web-curated image-text data created by the people, for the people

NeurIPS 2021poster

Large datasets of paired images and text have become increasingly popular for learning generic representations for vision and vision-and-language tasks. Such datasets have been built by querying search engines or collecting HTML alt-text – since web data is noisy, they require complex filtering pipe…

Cited by 183SourcecodeScholar
2019

Probabilistic Neural Symbolic Models for Interpretable Visual Question Answering

ICML 2019oral

We propose a new class of probabilistic neural-symbolic models, that have symbolic functional programs as a latent, stochastic variable. Instantiated in the context of visual question answering, our probabilistic formulation offers two key conceptual advantages over prior neural-symbolic models for…

Cited by 109SourcePDFScholar
2019

nocaps: novel object captioning at scale

ICCV 2019poster

Image captioning models have achieved impressive results on datasets containing limited visual concepts and large amounts of paired image-caption training data. However, if these models are to ever function in the wild, a much larger variety of visual concepts must be learned, ideally from less supe…

Cited by 420PDFcodeScholar