← Search

Martin J. Wainwright

12 accepted papers

2024

Taming "data-hungry" reinforcement learning? Stability in continuous state-action spaces

NeurIPS 2024poster

We introduce a novel framework for analyzing reinforcement learning (RL) in continuous state-action spaces, and use it to prove fast rates of convergence in both off-line and on-line settings. Our analysis highlights two key stability properties, relating to how changes in value functions and/or pol…

Cited by 4SourcePDFScholar
2020

Preference learning along multiple criteria: A game-theoretic perspective

NeurIPS 2020poster

The literature on ranking from ordinal data is vast, and there are several ways to aggregate overall preferences from pairwise comparisons between objects. In particular, it is well-known that any Nash equilibrium of the zero-sum game induced by the preference matrix defines a natural solution conce…

Cited by 16SourcePDFScholar
2019

L-Shapley and C-Shapley: Efficient Model Interpretation for Structured Data

ICLR 2019poster

Instancewise feature scoring is a method for model interpretation, which yields, for each test instance, a vector of importance scores associated with features. Methods based on the Shapley score have been proposed as a fair way of computing feature attributions, but incur an exponential complexity…

2018

Theoretical guarantees for EM under misspecified Gaussian mixture models

NeurIPS 2018poster

Recent years have witnessed substantial progress in understanding the behavior of EM for mixture models that are correctly specified. Given that model misspecification is common in practice, it is important to understand EM in this more general setting. We provide non-asymptotic guarantees…

Cited by 19SourcePDFScholar
2017

A framework for Multi-A(rmed)/B(andit) Testing with Online FDR Control

NeurIPS 2017spotlight

We propose an alternative framework to existing setups for controlling false alarms when multiple A/B tests are run over time. This setup arises in many practical applications, e.g. when pharmaceutical companies test new treatment options against control pills for different diseases, or when interne…

2017

Early stopping for kernel boosting algorithms: A general analysis with localized complexities

NeurIPS 2017spotlight

Early stopping of iterative algorithms is a widely-used form of regularization in statistical learning, commonly used in conjunction with boosting and related gradient-type algorithms. Although consistency results have been established in some settings, such estimators are less well-understood than…

Cited by 71SourcePDFScholar
2017

Kernel Feature Selection via Conditional Covariance Minimization

NeurIPS 2017poster

We propose a method for feature selection that employs kernel-based measures of independence to find a subset of covariates that is maximally predictive of the response. Building on past work in kernel dimension reduction, we show how to perform feature selection via a constrained optimization probl…

2017

Online control of the false discovery rate with decaying memory

NeurIPS 2017oral

In the online multiple testing problem, p-values corresponding to different null hypotheses are presented one by one, and the decision of whether to reject a hypothesis must be made immediately, after which the next p-value is presented. Alpha-investing algorithms to control the false discovery rate…

Cited by 80SourcePDFScholar
2016

Local Maxima in the Likelihood of Gaussian Mixture Models: Structural Results and Algorithmic Consequences

NeurIPS 2016poster

We provide two fundamental results on the population (infinite-sample) likelihood function of Gaussian mixture models with $M \geq 3$ components. Our first main result shows that the population likelihood function has bad local maxima even in the special case of equally-weighted mixtures of well-sep…

Cited by 198SourcePDFScholar