← Search

Mihir Prabhudesai

9 accepted papers

2026

Solving Physics Olympiad via Reinforcement Learning on Physics Simulators

ICML 2026poster

We have witnessed remarkable advances in LLM reasoning capabilities with the advent of DeepSeek-R1. However, much of this progress has been fueled by the abundance of internet question–answer (QA) pairs—a major bottleneck going forward, since such data is limited in scale and concentrated mainly in …

Cited by 0SourceScholar
2025

Diffusion Beats Autoregressive in Data-Constrained Settings

NeurIPS 2025poster

Autoregressive (AR) models have long dominated the landscape of large language models, driving progress across a wide range of tasks. Recently, diffusion-based language models have emerged as a promising alternative, though their advantages over AR models remain underexplored. In this paper, we syst…

Cited by 0SourcecodeScholar
2023

Diffusion-TTA: Test-time Adaptation of Discriminative Models via Generative Feedback

NeurIPS 2023poster

The advancements in generative modeling, particularly the advent of diffusion models, have sparked a fundamental question: how can these models be effectively used for discriminative tasks? In this work, we find that generative models can be great test-time adapters for discriminative models. Our me…

2023

Test-time Adaptation with Slot-Centric Models

ICML 2023poster

Current visual detectors, though impressive within their training distribution, often fail to parse out-of-distribution scenes into their constituent entities. Recent test-time adaptation methods use auxiliary self-supervised losses to adapt the network parameters to each test example independently…

2023

Your Diffusion Model is Secretly a Zero-Shot Classifier

ICCV 2023poster

The recent wave of large-scale text-to-image diffusion models has dramatically increased our text-based image generation abilities. These models can generate realistic images for a staggering variety of prompts and exhibit impressive compositional generalization abilities. Almost all use cases thus…

Cited by 267PDFcodeScholar
2021

CoCoNets: Continuous Contrastive 3D Scene Representations

CVPR 2021poster

This paper explores self-supervised learning of amodal 3D feature representations from RGB and RGB-D posed images and videos, agnostic to object and scene semantic content, and evaluates the resulting scene representations in the downstream tasks of visual correspondence, object tracking, and object…

Cited by 27PDFScholar
2021

Disentangling 3D Prototypical Networks for Few-Shot Concept Learning

ICLR 2021poster

We present neural architectures that disentangle RGB-D images into objects’ shapes and styles and a map of the background scene, and explore their applications for few-shot 3D object detection and few-shot concept classification. Our networks incorporate architectural biases that reflect the image f…

2020

3D-OES: Viewpoint-Invariant Object-Factorized Environment Simulators

CoRL 2020

We propose an action-conditioned dynamics model that predicts scene changes caused by object and agent interactions in a viewpoint-invariant 3D neural scene representation space, inferred from RGB-D videos. In this 3D feature space, objects do not interfere with one another and their appearance pers

Cited by 0SourcePDFScholar
2020

Embodied Language Grounding With 3D Visual Feature Representations

CVPR 2020poster

We propose associating language utterances to 3D visual abstractions of the scene they describe. The 3D visual abstractions are encoded as 3-dimensional visual feature maps. We infer these 3D visual scene feature maps from RGB images of the scene via view prediction: when the generated 3D scene feat…

Cited by 23PDFScholar