← Search

Venkatesh Saligrama

69 accepted papers

2026

BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models

CVPR 2026

Early children's developmental trajectories set up a natural goal for sample-efficient pretraining of vision foundation models. We introduce BabyVLM-V2, a developmentally grounded framework for infant-inspired vision-language modeling that extensively improves upon BabyVLM-V1 through a longitudinal,

Cited by 0SourcecodeScholar
2026

Symmetry Reveals the In-Context Classifier: Transformers Implement Mean-Shift Dynamics

ICML 2026spotlight

Transformers can perform in-context classification from a few labeled examples, yet the inference-time algorithm remains opaque. We study multi-class linear classification in the hard no-margin regime and make the computation identifiable by enforcing feature- and label-permutation equivariance at e…

Cited by 0SourceScholar
2025

BabyVLM: Data-Efficient Pretraining of VLMs Inspired by Infant Learning

ICCV 2025poster

Human infants rapidly develop visual reasoning skills from minimal input, suggesting that developmentally inspired pretraining could significantly enhance the efficiency of vision-language models (VLMs). Although recent efforts have leveraged infant-inspired datasets like SAYCam, existing evaluation…

Cited by 0SourcePDFScholar
2025

Feasible Action Search for Bandit Linear Programs via Thompson Sampling

ICML 2025poster

We study the 'feasible action search' (FAS) problem for linear bandits, wherein a learner attempts to discover a feasible point for a set of linear constraints $\Phi_* a \ge 0,$ without knowledge of the matrix $\Phi_* \in \mathbb{R}^{m \times d}$. A FAS learner selects a sequence of actions $a_t,$ a…

Cited by 0SourcePDFScholar
2025

GPS: A Probabilistic Distributional Similarity with Gumbel Priors for Set-to-Set Matching

ICLR 2025poster

Set-to-set matching aims to identify correspondences between two sets of unordered items by minimizing a distance metric or maximizing a similarity measure. Traditional metrics, such as Chamfer Distance (CD) and Earth Mover’s Distance (EMD), are widely used for this purpose but often suffer from lim…

2025

Linear Transformers Implicitly Discover Unified Numerical Algorithms

NeurIPS 2025poster

A transformer is merely a stack of learned data–to–data maps—yet those maps can hide rich algorithms. We train a linear, attention-only transformer on millions of masked-block completion tasks: each prompt is a masked low-rank matrix whose missing block may be (i) a scalar prediction target or (ii)…

Cited by 0SourceScholar
2025

SPARC: Score Prompting and Adaptive Fusion for Zero-Shot Multi-Label Recognition in Vision-Language Models

CVPR 2025poster

Zero-shot multi-label recognition (MLR) with Vision-Language Models (VLMs) faces significant challenges without training data, model tuning, or architectural modifications. Existing approaches require prompt tuning or architectural adaptations, limiting zero-shot applicability. Our work proposes a n…

2025

Scaling Up Temporal Domain Generalization via Temporal Experts Averaging

EMNLP 2025

Temporal Domain Generalization (TDG) aims to generalize across temporal distribution shifts, e.g., lexical change over time. Prior work often addresses this by predicting future model weights. However, full model prediction is prohibitively expensive for even reasonably sized models. Thus, recent me

2024

Testing the Feasibility of Linear Programs with Bandit Feedback

ICML 2024spotlight

While the recent literature has seen a surge in the study of constrained bandit problems, all existing methods for these begin by assuming the feasibility of the underlying problem. We initiate the study of testing such feasibility assumptions, and in particular address the problem in the linear ban…

Cited by 0SourcePDFScholar
2023

Efficient Edge Inference by Selective Query

ICLR 2023poster

Edge devices provide inference on predictive tasks to many end-users. However, deploying deep neural networks that achieve state-of-the-art accuracy on these devices is infeasible due to edge resource constraints. Nevertheless, cloud-only processing, the de-facto standard, is also problematic, since…

Cited by 22SourcePDFScholar
2023

Ideology Prediction from Scarce and Biased Supervision: Learn to Disregard the “What” and Focus on the “How”!

ACL 2023long

We propose a novel supervised learning approach for political ideology prediction (PIP) that is capable of predicting out-of-distribution inputs. This problem is motivated by the fact that manual data-labeling is expensive, while self-reported labels are often scarce and exhibit significant selectio…

Cited by 5SourcePDFScholar
2023

InfoCD: A Contrastive Chamfer Distance Loss for Point Cloud Completion

NeurIPS 2023poster

A point cloud is a discrete set of data points sampled from a 3D geometric surface. Chamfer distance (CD) is a popular metric and training loss to measure the distances between point clouds, but also well known to be sensitive to outliers. To address this issue, in this paper we propose InfoCD, a no…

2023

Learning Human Action Recognition Representations Without Real Humans

NeurIPS 2023poster

Pre-training on massive video datasets has become essential to achieve high action recognition performance on smaller downstream datasets. However, most large-scale video datasets contain images of people and hence are accompanied with issues related to privacy, ethics, and data protection, often pr…

2022

Faster Algorithms for Learning Convex Functions

ICML 2022spotlight

The task of approximating an arbitrary convex function arises in several learning problems such as convex regression, learning with a difference of convex (DC) functions, and learning Bregman or $f$-divergences. In this paper, we develop and analyze an approach for solving a broad range of convex fu…

2022

How Transferable are Video Representations Based on Synthetic Data?

NeurIPS 2022accept

Action recognition has improved dramatically with massive-scale video datasets. Yet, these datasets are accompanied with issues related to curation cost, privacy, ethics, bias, and copyright. Compared to that, only minor efforts have been devoted toward exploring the potential of synthetic video dat…

2022

Strategies for Safe Multi-Armed Bandits with Logarithmic Regret and Risk

ICML 2022spotlight

We investigate a natural but surprisingly unstudied approach to the multi-armed bandit problem under safety risk constraints. Each arm is associated with an unknown law on safety risks and rewards, and the learner’s goal is to maximise reward whilst not playing unsafe arms, as determined by a given…

Cited by 14SourcePDFScholar
2022

Task2Sim: Towards Effective Pre-Training and Transfer From Synthetic Data

CVPR 2022poster

Pre-training models on Imagenet or other massive datasets of real images has led to major advances in computer vision, albeit accompanied with shortcomings related to curation cost, privacy, usage rights, and ethical issues. In this paper, for the first time, we study the transferability of pre-trai…

Cited by 47PDFScholar
2021

Debiasing Model Updates for Improving Personalized Federated Training

ICML 2021spotlight

We propose a novel method for federated learning that is customized specifically to the objective of a given edge device. In our proposed method, a server trains a global meta-model by collaborating with devices without actually sharing data. The trained global meta-model is then personalized locall…

Cited by 86SourcePDFScholar
2021

Effectively Leveraging Attributes for Visual Similarity

ICCV 2021poster

Measuring similarity between two images often requires performing complex reasoning along different axes (e.g., color, texture, or shape). Insights into what might be important for measuring similarity can can be provided by annotated attributes, but prior work tends to view these annotations as com…

Cited by 13PDFcodeScholar
2021

Federated Learning Based on Dynamic Regularization

ICLR 2021oral

We propose a novel federated learning method for distributively training neural network models, where the server orchestrates cooperation between a subset of randomly chosen devices in each round. We view Federated Learning problem primarily from a communication perspective and allow more device lev…

2021

Online Selective Classification with Limited Feedback

NeurIPS 2021spotlight

Motivated by applications to resource-limited and safety-critical domains, we study selective classification in the online learning model, wherein a predictor may abstain from classifying an instance. For example, this may model an adaptive decision to invoke more resources on this instance. Two sal…

2021

Training Recurrent Neural Networks via Forward Propagation Through Time

ICML 2021spotlight

Back-propagation through time (BPTT) has been widely used for training Recurrent Neural Networks (RNNs). BPTT updates RNN parameters on an instance by back-propagating the error in time over the entire sequence length, and as a result, leads to poor trainability due to the well-known gradient explos…

2020

Learning to Approximate a Bregman Divergence

NeurIPS 2020poster

Bregman divergences generalize measures such as the squared Euclidean distance and the KL divergence, and arise throughout many areas of machine learning. In this paper, we focus on the problem of approximating an arbitrary Bregman divergence from supervision, and we provide a well-principled appro…

2020

Minimax Rate for Learning From Pairwise Comparisons in the BTL Model

ICML 2020poster

We consider the problem of learning the qualities w_1, ... , w_n of a collection of items by performing noisy comparisons among them. We assume there is a fixed “comparison graph” and every neighboring pair of items is compared k times. We will study the popular Bradley-Terry-Luce model, where the p…

Cited by 24SourcePDFScholar
2020

Online Algorithm for Unsupervised Sequential Selection with Contextual Information

NeurIPS 2020poster

In this paper, we study Contextual Unsupervised Sequential Selection (USS), a new variant of the stochastic contextual bandits problem where the loss of an arm cannot be inferred from the observed feedback. In our setup, arms are associated with fixed costs and are ordered, forming a cascade. In eac…

Cited by 5SourcePDFScholar
2020

Piecewise Linear Regression via a Difference of Convex Functions

ICML 2020poster

We present a new piecewise linear regression methodology that utilises fitting a \emph{difference of convex} functions (DC functions) to the data. These are functions $f$ that may be represented as the difference $\phi_1 - \phi_2$ for a choice of \emph{convex} functions $\phi_1, \phi_2$. The method…

2020

RNNs Incrementally Evolving on an Equilibrium Manifold: A Panacea for Vanishing and Exploding Gradients?

ICLR 2020poster

Recurrent neural networks (RNNs) are particularly well-suited for modeling long-term dependencies in sequential data, but are notoriously hard to train because the error backpropagated in time either vanishes or explodes at an exponential rate. While a number of works attempt to mitigate this effect…

Cited by 68SourcecodeScholar
2019

Cost aware Inference for IoT Devices

AISTATS 2019poster

Networked embedded devices (IoTs) of limited CPU, memory and power resources are revolutionizing data gathering, remote monitoring and planning in many consumer and business applications. Nevertheless, resource limitations place a significant burden on their service life and operation, warranting co…

Cited by 15SourcePDFScholar
2019

Cost-Aware Fine-Grained Recognition for IoTs Based on Sequential Fixations

ICCV 2019poster

We consider the problem of fine-grained classification on an edge camera device that has limited power. The edge device must sparingly interact with the cloud to minimize communication bits to conserve power, and the cloud upon receiving the edge inputs returns a classification label. To deal with f…

Cited by 3PDFScholar
2019

Efficient Near-Optimal Testing of Community Changes in Balanced Stochastic Block Models

NeurIPS 2019poster

We propose and analyze the problems of \textit{community goodness-of-fit and two-sample testing} for stochastic block models (SBM), where changes arise due to modification in community memberships of nodes. Motivated by practical applications, we consider the challenging sparse regime, where expect…

Cited by 9SourcePDFScholar
2019

Graph Resistance and Learning from Pairwise Comparisons

ICML 2019oral

We consider the problem of learning the qualities of a collection of items by performing noisy comparisons among them. Following the standard paradigm, we assume there is a fixed “comparison graph” and every neighboring pair of items in this graph is compared k times according to the Bradley-Terry-L…

Cited by 15SourcePDFScholar
2019

Learning Classifiers for Target Domain with Limited or No Labels

ICML 2019oral

In computer vision applications, such as domain adaptation (DA), few shot learning (FSL) and zero-shot learning (ZSL), we encounter new objects and environments, for which insufficient examples exist to allow for training “models from scratch,” and methods that adapt existing models, trained on the…

Cited by 15SourcePDFScholar
2019

Online Algorithm for Unsupervised Sensor Selection

AISTATS 2019poster

In many security and healthcare systems, the detection and diagnosis systems use a sequence of sensors/tests. Each test outputs a prediction of the latent state and carries an inherent cost. However, the correctness of the predictions cannot be evaluated due to unavailability of the ground-truth ann…

Cited by 14SourcePDFScholar
2019

Shallow RNN: Accurate Time-series Classification on Resource Constrained Devices

NeurIPS 2019poster

Recurrent Neural Networks (RNNs) capture long dependencies and context, and 2 hence are the key component of typical sequential data based tasks. However, the sequential nature of RNNs dictates a large inference cost for long sequences even if the hardware supports parallelization. To induce long-te…

2018

Gradient Descent for Sparse Rank-One Matrix Completion for Crowd-Sourced Aggregation of Sparsely Interacting Workers

ICML 2018oral

We consider worker skill estimation for the single coin Dawid-Skene crowdsourcing model. In practice skill-estimation is challenging because worker assignments are sparse and irregular due to the arbitrary, and uncontrolled availability of workers. We formulate skill estimation as a rank-one correla…

Cited by 31SourcePDFScholar
2018

Two-Sample Testing can be as Hard as Structure Learning in Ising Models: Minimax Lower Bounds

ICASSP 2018accepted

Consider the following structural two-sample testing problem: given two sets of sample drawn from Ising models, determine whether the underlying network structure has changed. In [1], we showed that for Ising models over p variables with network structures that have degree bounded by d, under mild c…

Cited by 0SourceScholar
2016

Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings

NeurIPS 2016poster

The blind application of machine learning runs the risk of amplifying biases present in data. Such a danger is facing us with word embedding, a popular framework to represent text data as vectors which has been used in many machine learning and natural language processing tasks. We show that even wo…

Cited by 4470SourcePDFScholar
2015

Efficient Learning by Directed Acyclic Graph For Resource Constrained Prediction

NeurIPS 2015poster

We study the problem of reducing test-time acquisition costs in classification systems. Our goal is to learn decision rules that adaptively select sensors for each example as necessary to make a confident prediction. We model our system as a directed acyclic graph (DAG) where internal nodes correspo…

Cited by 71SourcePDFScholar
2015

Learning shared rankings from mixtures of noisy pairwise comparisons

ICASSP 2015accepted

We propose a novel model for rank aggregation from pairwise comparisons which accounts for a heterogeneous population of inconsistent users whose preferences are different mixtures of multiple shared ranking schemes. By connecting this problem to recent advances in the non-negative matrix factorizat…

Cited by 0SourceScholar