← Search

Arkabandhu Chowdhury

4 accepted papers

2024

CALVIN: Improved Contextual Video Captioning via Instruction Tuning

NeurIPS 2024poster

The recent emergence of powerful Vision-Language models (VLMs) has significantly improved image captioning. Some of these models are extended to caption videos as well. However, their capabilities to understand complex scenes are limited, and the descriptions they provide for scenes tend to be overl…

Cited by 0SourcePDFScholar
2023

Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles

ICML 2023oral

Modern hierarchical vision transformers have added several vision-specific components in the pursuit of supervised classification performance. While these components lead to effective accuracies and attractive FLOP counts, the added complexity actually makes these transformers slower than their vani…

2021

Few-Shot Image Classification: Just Use a Library of Pre-Trained Feature Extractors and a Simple Classifier

ICCV 2021poster

Recent papers have suggested that transfer learning can outperform sophisticated meta-learning methods for few-shot image classification. We take this hypothesis to its logical conclusion, and suggest the use of an ensemble of high-quality, pre-trained feature extractors for few-shot image classific…

Cited by 45PDFcodeScholar