← Search

Ankit Singh

7 accepted papers

2026

SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models

CVPR 2026

Vision foundation models trained via multi-teacher distillation offer a promising path toward unified visual representations, yet the learning dynamics and data efficiency of such approaches remain underexplored. In this paper, we systematically study multi-teacher distillation for vision foundation

Cited by 0SourcecodeScholar
2026

VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs

CVPR 2026

Vision-Language Models (VLMs) have achieved remarkable progress across tasks such as visual question answering and image captioning. Yet, the extent to which these models perform visual reasoning as opposed to relying on linguistic priors remains unclear. To address this, we introduce VisRes Bench,

Cited by 0SourcecodeScholar
2025

Harnessing Frozen Unimodal Encoders for Flexible Multimodal Alignment

CVPR 2025poster

Recent contrastive multimodal vision-language models like CLIP have demonstrated robust open-world semantic understanding, becoming the standard image backbones for vision-language applications. However, recent findings suggest high semantic similarity between well-trained unimodal encoders, which r…

2025

Vision-Language Models Can't See the Obvious

ICCV 2025poster

We present Saliency Benchmark (SalBench), a novel benchmark designed to assess the capability of Large Vision-Language Models (LVLM) in detecting visually salient features that are readily apparent to humans, such as a large circle amidst a grid of smaller ones. This benchmark focuses on low-level f…

Cited by 0SourcePDFScholar
2023

On permutation symmetries in Bayesian neural network posteriors: a variational perspective

NeurIPS 2023poster

The elusive nature of gradient-based optimization in neural networks is tied to their loss landscape geometry, which is poorly understood. However recent work has brought solid evidence that there is essentially no loss barrier between the local solutions of gradient descent, once accounting for wei…

Cited by 5SourcePDFScholar
2021

Semi-Supervised Action Recognition With Temporal Contrastive Learning

CVPR 2021poster

Learning to recognize actions from only a handful of labeled videos is a challenging problem due to the scarcity of tediously collected activity labels. We approach this problem by learning a two-pathway temporal contrastive model using unlabeled videos at two different speeds leveraging the fact th…

Cited by 134PDFcodeScholar