← Search

Houssam Nassif

7 accepted papers

2026

Fixed Budget is No Harder Than Fixed Confidence in Best-Arm Identification up to Logarithmic Factors

ICML 2026poster

The best-arm identification (BAI) problem is one of the most fundamental problems in interactive machine learning, which has two flavors: the fixed-budget setting (FB) and the fixed-confidence setting (FC). For $K$-armed bandits with the unique best arm, the optimal sample complexities for both sett…

Cited by 0SourceScholar
2024

On Neural Networks as Infinite Tree-Structured Probabilistic Graphical Models

NeurIPS 2024poster

Deep neural networks (DNNs) lack the precise semantics and definitive probabilistic interpretation of probabilistic graphical models (PGMs). In this paper, we propose an innovative solution by constructing infinite tree-structured PGMs that correspond exactly to neural networks. Our research reveals…

2023

A Data-Driven State Aggregation Approach for Dynamic Discrete Choice Models

UAI 2023poster

In dynamic discrete choice models, a commonly studied problem is estimating parameters of agent reward functions (also known as ’structural’ parameters) using agent behavioral data. This task is also known as inverse reinforcement learning. Maximum likelihood estimation for such models requires dyna…

2023

Experimental Designs for Heteroskedastic Variance

NeurIPS 2023poster

Most linear experimental design problems assume homogeneous variance, while the presence of heteroskedastic noise is present in many realistic settings. Let a learner have access to a finite set of measurement vectors $\mathcal{X}\subset \mathbb{R}^d$ that can be probed to receive noisy linear resp…

Cited by 5SourcePDFScholar
2022

Instance-optimal PAC Algorithms for Contextual Bandits

NeurIPS 2022accept

In the stochastic contextual bandit setting, regret-minimizing algorithms have been extensively researched, but their instance-minimizing best-arm identification counterparts remain seldom studied. In this work, we focus on the stochastic bandit problem in the $(\epsilon,\delta)$-PAC setting: given…

Cited by 29SourcePDFScholar
2021

Improved Confidence Bounds for the Linear Logistic Model and Applications to Bandits

ICML 2021spotlight

We propose improved fixed-design confidence bounds for the linear logistic model. Our bounds significantly improve upon the state-of-the-art bound by Li et al. (2017) via recent developments of the self-concordant analysis of the logistic loss (Faury et al., 2020). Specifically, our confidence bound…

Cited by 29SourcePDFScholar
2020

Deep PQR: Solving Inverse Reinforcement Learning using Anchor Actions

ICML 2020accepted

We propose a reward function estimation framework for inverse reinforcement learning with deep energy-based policies. We name our method PQR, as it sequentially estimates the Policy, the Q-function, and the Reward function by deep learning. PQR does not assume that the reward solely depends on the s…

Cited by 22SourcePDFScholar