← Search

Miroslav Dudík

6 accepted papers

2024

PcLast: Discovering Plannable Continuous Latent States

ICML 2024poster

Goal-conditioned planning benefits from learned low-dimensional representations of rich observations. While compact latent representations typically learned from variational autoencoders or inverse dynamics enable goal-conditioned decision making, they ignore state reachability, hampering their perf…

Cited by 2SourcePDFScholar
2024

SureMap: Simultaneous mean estimation for single-task and multi-task disaggregated evaluation

NeurIPS 2024poster

Disaggregated evaluation—estimation of performance of a machine learning model on different subpopulations—is a core task when assessing performance and group-fairness of AI systems. A key challenge is that evaluation data is scarce, and subpopulations arising from intersections of attri…

2023

A Unified Model and Dimension for Interactive Estimation

NeurIPS 2023poster

We study an abstract framework for interactive learning called interactive estimation in which the goal is to estimate a target from its ``similarity'' to points queried by the learner. We introduce a combinatorial measure called Dissimilarity dimension which largely captures learnability in our mod…

Cited by 0SourcePDFScholar
2022

Provably sample-efficient RL with side information about latent dynamics

NeurIPS 2022accept

We study reinforcement learning (RL) in settings where observations are high-dimensional, but where an RL agent has access to abstract knowledge about the structure of the state space, as is the case, for example, when a robot is tasked to go to a specific room in a building using observations from…

Cited by 2SourcePDFScholar
2021

Bayesian decision-making under misspecified priors with applications to meta-learning

NeurIPS 2021spotlight

Thompson sampling and other Bayesian sequential decision-making algorithms are among the most popular approaches to tackle explore/exploit trade-offs in (contextual) bandits. The choice of prior in these algorithms offers flexibility to encode domain knowledge but can also lead to poor performance…

Cited by 62SourcePDFScholar