← Search

Parnian Kassraie

9 accepted papers

2025

LITE: Efficiently Estimating Gaussian Probability of Maximality

AISTATS 2025poster

We consider the problem of computing the *probability of maximality* (PoM) of a Gaussian random vector, i.e., the probability for each dimension to be maximal. This is a key challenge in applications ranging from Bayesian optimization to reinforcement learning, where the PoM not only helps with find…

Cited by 0SourcecodeScholar
2024

Bandits with Preference Feedback: A Stackelberg Game Perspective

NeurIPS 2024poster

Bandits with preference feedback present a powerful tool for optimizing unknown target functions when only pairwise comparisons are allowed instead of direct value queries. This model allows for incorporating human feedback into online inference and optimization and has been employed in systems for…

2024

Progressive Entropic Optimal Transport Solvers

NeurIPS 2024poster

Optimal transport (OT) has profoundly impacted machine learning by providing theoretical and computational tools to realign datasets. In this context, given two large point clouds of sizes $n$ and $m$ in $\mathbb{R}^d$, entropic OT (EOT) solvers have emerged as the most reliable tool to either solve…

Cited by 4SourcePDFScholar
2023

Anytime Model Selection in Linear Bandits

NeurIPS 2023poster

Model selection in the context of bandit optimization is a challenging problem, as it requires balancing exploration and exploitation not only for action selection, but also for model selection. One natural approach is to rely on online learning algorithms that treat different models as experts. Exi…

2023

Hallucinated adversarial control for conservative offline policy evaluation

UAI 2023poster

We study the problem of conservative off-policy evaluation (COPE) where given an offline dataset of environment interactions, collected by other agents, we seek to obtain a (tight) lower bound on a policy’s performance. This is crucial when deciding whether a given policy satisfies certain minimal p…

2023

Lifelong bandit optimization: no prior and no regret

UAI 2023poster

Machine learning algorithms are often repeatedly. applied to problems with similar structure over and over again. We focus on solving a sequence of bandit optimization tasks and develop LIBO, an algorithm which adapts to the environment by learning from past experience and becomes more sample-effici…

Cited by 7SourcePDFScholar
2022

Meta-Learning Hypothesis Spaces for Sequential Decision-making

ICML 2022spotlight

Obtaining reliable, adaptive confidence sets for prediction functions (hypotheses) is a central challenge in sequential decision-making tasks, such as bandits and model-based reinforcement learning. These confidence sets typically rely on prior assumptions on the hypothesis space, e.g., the known ke…

Cited by 10SourcePDFScholar