← Search

Sharut Gupta

10 accepted papers

2026

Better Together: Leveraging Unpaired Multimodal Data for Stronger Unimodal Models

ICLR 2026poster

Traditional multimodal learners find unified representations for tasks like visual question answering, but rely heavily on large paired datasets. However, an overlooked yet potentially powerful question is: can one leverage auxiliary $\textit{unpaired}$ multimodal data to directly enhance representa…

Cited by 0SourcecodeScholar
2026

Sequential Parallel Duality in Prefix Scannable Models

ICLR 2026poster

Modern neural sequence models are designed to meet the dual mandate of parallelizable training and fast sequential inference. Recent developments have given rise to various models, such as Gated Linear Attention (GLA) and Mamba, that achieve such ``sequential-parallel duality.'' This raises a natura…

Cited by 0SourceScholar
2025

An Information Criterion for Controlled Disentanglement of Multimodal Data

ICLR 2025poster

Multimodal representation learning seeks to relate and decompose information inherent in multiple modalities. By disentangling modality-specific information from information that is shared across modalities, we can improve interpretability and robustness and enable downstream tasks such as the gener…

2025

Learning Diffusion Models with Flexible Representation Guidance

NeurIPS 2025poster

Diffusion models can be improved with additional guidance towards more effective representations of input. Indeed, prior empirical work has already shown that aligning internal representations of the diffusion model with those of pre-trained models improves generation quality. In this paper, we pres…

Cited by 0SourceScholar
2024

In-Context Symmetries: Self-Supervised Learning through Contextual World Models

NeurIPS 2024poster

At the core of self-supervised learning for vision is the idea of learning invariant or equivariant representations with respect to a set of data transformations. This approach, however, introduces strong inductive biases, which can render the representations fragile in downstream tasks that do not…

Cited by 0SourcePDFScholar
2024

Removing Biases from Molecular Representations via Information Maximization

ICLR 2024poster

High-throughput drug screening -- using cell imaging or gene expression measurements as readouts of drug effect -- is a critical tool in biotechnology to assess and understand the relationship between the chemical structure and biological activity of a drug. Since large-scale screens have to be divi…

2024

Structuring Representation Geometry with Rotationally Equivariant Contrastive Learning

ICLR 2024poster

Self-supervised learning converts raw perceptual data such as images to a compact space where simple Euclidean distances measure meaningful variations in data. In this paper, we extend this formulation by adding additional geometric structure to the embedding space by enforcing transformations of in…

2024

Understanding the Role of Equivariance in Self-supervised Learning

NeurIPS 2024poster

Contrastive learning has been a leading paradigm for self-supervised learning, but it is widely observed that it comes at the price of sacrificing useful features (\eg colors) by being invariant to data augmentations. Given this limitation, there has been a surge of interest in equivariant self-supe…

2022

AdaBest: Minimizing Client Drift in Federated Learning via Adaptive Bias Estimation

ECCV 2022poster

"In Federated Learning (FL), a number of clients or devices collaborate to train a model without sharing their data. Models are optimized locally at each client and further communicated to a central hub for aggregation. While FL is an appealing decentralized training paradigm, heterogeneity among da…

Cited by 26SourcePDFScholar