← Search

Shobhita Sundaram

5 accepted papers

2026

Better Together: Leveraging Unpaired Multimodal Data for Stronger Unimodal Models

ICLR 2026poster

Traditional multimodal learners find unified representations for tasks like visual question answering, but rely heavily on large paired datasets. However, an overlooked yet potentially powerful question is: can one leverage auxiliary $\textit{unpaired}$ multimodal data to directly enhance representa…

Cited by 0SourcecodeScholar
2026

Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability

ICML 2026spotlight

RL methods for finetuning large reasoning models stall on datasets with low initial success rates, and thus little training signal. We investigate a fundamental question: Can a pretrained LLM leverage latent knowledge to generate an automated curriculum for problems it cannot solve? We explore this …

Cited by 0SourceScholar
2025

Personalized Representation from Personalized Generation

ICLR 2025poster

Modern vision models excel at general purpose downstream tasks. It is unclear, however, how they may be used for personalized vision tasks, which are both fine-grained and data-scarce. Recent works have successfully applied synthetic data to general-purpose representation learning, while advances in…

2024

When does perceptual alignment benefit vision representations?

NeurIPS 2024poster

Humans judge perceptual similarity according to diverse visual attributes, including scene layout, subject location, and camera pose. Existing vision models understand a wide range of semantic abstractions but improperly weigh these attributes and thus make inferences misaligned with human perceptio…

Cited by 5SourcePDFScholar
2023

DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data

NeurIPS 2023spotlight

Current perceptual similarity metrics operate at the level of pixels and patches. These metrics compare images in terms of their low-level colors and textures, but fail to capture mid-level similarities and differences in image layout, object pose, and semantic content. In this paper, we develop a p…