← Search

Vinay Kumar Verma

12 accepted papers

2026

FOCUS: Forcing In-Context Object Localization through Visual Support Constraints and Policy Optimization

ICML 2026poster

In-context localization (ICL) seeks to localize a target object specified by a small set of support examples in a query image, operating on the fly without training or parameter updates. Despite rapid advances in vision–language models (VLMs), achieving category-agnostic and visually grounded ICL re…

Cited by 0SourceScholar
2026

TaTToo: Tool-Grounded Thinking PRM for Test-Time Scaling in Tabular Reasoning

ICLR 2026poster

Process Reward Models (PRMs) have recently emerged as a powerful framework for enhancing the reasoning capabilities of large reasoning models (LRMs), particularly in the context of test-time scaling (TTS). However, their potential for supervising LRMs on tabular reasoning domains remains underexplor…

Cited by 0SourceScholar
2025

MoEMoE: Question Guided Dense and Scalable Sparse Mixture-of-Expert for Multi-source Multi-modal Answering

NAACL 2025industry

Question Answering (QA) and Visual Question Answering (VQA) are well-studied problems in the language and vision domain. One challenging scenario involves multiple sources of information, each of a different modality, where the answer to the question may exist in one or more sources. This scenario c…

Cited by 0SourcePDFScholar
2025

VADE: Visual Attention Guided Hallucination Detection and Elimination

ACL 2025finding

Vision Language Models (VLMs) have achieved significant advancements in complex visual understanding tasks. However, VLMs are prone to hallucinations—generating outputs that lack alignment with visual content. This paper addresses hallucination detection in VLMs by leveraging the visual grounding in…

Cited by 0SourcePDFScholar
2021

CAM-GAN: Continual Adaptation Modules for Generative Adversarial Networks

NeurIPS 2021poster

We present a continual learning approach for generative adversarial networks (GANs), by designing and leveraging parameter-efficient feature map transformations. Our approach is based on learning a set of global and task-specific parameters. The global parameters are fixed across tasks whereas the t…

Cited by 32SourcePDFScholar
2021

Continual Learning using a Bayesian Nonparametric Dictionary of Weight Factors

AISTATS 2021poster

Naively trained neural networks tend to experience catastrophic forgetting in sequential task settings, where data from previous tasks are unavailable. A number of methods, using various model expansion strategies, have been proposed recently as possible solutions. However, determining how much to e…

Cited by 41SourcePDFScholar
2021

Efficient Feature Transformations for Discriminative and Generative Continual Learning

CVPR 2021poster

As neural networks are increasingly being applied to real-world applications, mechanisms to address distributional shift and sequential task learning without forgetting are critical. Methods incorporating network expansion have shown promise by naturally adding model capacity for learning new tasks…

Cited by 90PDFcodeScholar
2021

Knowledge Consolidation based Class Incremental Online Learning with Limited Data

IJCAI 2021poster

We propose a novel approach for class incremental online learning in a limited data setting. This problem setting is challenging because of the following constraints: (1) Classes are given incrementally, which necessitates a class incremental learning approach; (2) Data for each class is given in a…

Cited by 0SourcePDFScholar
2020

Calibrating CNNs for Lifelong Learning

NeurIPS 2020poster

We present an approach for lifelong/continual learning of convolutional neural networks (CNN) that does not suffer from the problem of catastrophic forgetting when moving from one task to the other. We show that the activation maps generated by the CNN trained on the old task can be calibrated using…

2019

HetConv: Heterogeneous Kernel-Based Convolutions for Deep CNNs

CVPR 2019poster

We present a novel deep learning architecture in which the convolution operation leverages heterogeneous kernels. The proposed HetConv (Heterogeneous Kernel-Based Convolution) reduces the computation (FLOPs) and the number of parameters as compared to standard convolution operation while still maint…

Cited by 144PDFScholar
2018

Generalized Zero-Shot Learning via Synthesized Examples

CVPR 2018poster

We present a generative framework for generalized zero-shot learning where the training and test classes are not necessarily disjoint. Built upon a variational autoencoder based architecture, consisting of a probabilistic encoder and a probabilistic emph{conditional} decoder, our model can generate…

Cited by 570SourcePDFScholar