← Search

Jyoti Aneja

5 accepted papers

2022

ELEVATER: A Benchmark and Toolkit for Evaluating Language-Augmented Visual Models

NeurIPS 2022accept

Learning visual representations from natural language supervision has recently shown great promise in a number of pioneering works. In general, these language-augmented visual models demonstrate strong transferability to a variety of datasets/tasks. However, it remains challenging to evaluate the tr…

Cited by 159SourcePDFScholar
2021

A Contrastive Learning Approach for Training Variational Autoencoder Priors

NeurIPS 2021poster

Variational autoencoders (VAEs) are one of the powerful likelihood-based generative models with applications in many domains. However, they struggle to generate high-quality images, especially when samples are obtained from the prior without any tempering. One explanation for VAEs' poor generative q…

Cited by 97SourcePDFScholar
2019

Fast, Diverse and Accurate Image Captioning Guided by Part-Of-Speech

CVPR 2019oral

Image captioning is an ambiguous problem, with many suitable captions for an image. To address ambiguity, beam search is the de facto method for sampling multiple captions. However, beam search is computationally expensive and known to produce generic captions. To address this concern, some vari…

Cited by 172PDFScholar
2019

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning

ICCV 2019poster

Diverse and accurate vision+language modeling is an important goal to retain creative freedom and maintain user engagement. However, adequately capturing the intricacies of diversity in language models is challenging. Recent works commonly resort to latent variable models augmented with more or less…

Cited by 84PDFScholar