← Search

Andras Gyorgy

12 accepted papers

2022

Faster Rates, Adaptive Algorithms, and Finite-Time Bounds for Linear Composition Optimization and Gradient TD Learning

AISTATS 2022poster

Gradient temporal difference (GTD) algorithms are provably convergent policy evaluation methods for off-policy reinforcement learning. Despite much progress, proper tuning of the stochastic approximation methods used to solve the resulting saddle point optimization problem requires the knowledge of…

Cited by 1SourcePDFScholar
2021

Confident Off-Policy Evaluation and Selection through Self-Normalized Importance Weighting

AISTATS 2021poster

We consider off-policy evaluation in the contextual bandit setting for the purpose of obtaining a robust off-policy selection strategy, where the selection strategy is evaluated based on the value of the chosen policy in a set of proposal (target) policies. We propose a new method to compute a lower…

2020

A FRAMEWORK FOR ROBUSTNESS CERTIFICATION OF SMOOTHED CLASSIFIERS USING F-DIVERGENCES

ICLR 2020poster

Formal verification techniques that compute provable guarantees on properties of machine learning models, like robustness to norm-bounded adversarial perturbations, have yielded impressive results. Although most techniques developed so far require knowledge of the architecture of the machine learnin…

Cited by 64SourceScholar
2020

A simpler approach to accelerated optimization: iterative averaging meets optimism

ICML 2020poster

Recently there have been several attempts to extend Nesterov’s accelerated algorithm to smooth stochastic and variance-reduced optimization. In this paper, we show that there is a simpler approach to acceleration: applying optimistic online learning algorithms and querying the gradient oracle at the…

Cited by 40SourcePDFScholar
2019

CapsAndRuns: An Improved Method for Approximately Optimal Algorithm Configuration

ICML 2019oral

We consider the problem of configuring general-purpose solvers to run efficiently on problem instances drawn from an unknown distribution, a problem of major interest in solver autoconfiguration. Following previous work, we focus on designing algorithms that find a configuration with near-optimal ex…

Cited by 27SourcePDFScholar
2019

Learning from Delayed Outcomes via Proxies with Applications to Recommender Systems

ICML 2019oral

Predicting delayed outcomes is an important problem in recommender systems (e.g., if customers will finish reading an ebook). We formalize the problem as an adversarial, delayed online learning problem and consider how a proxy for the delayed outcome (e.g., if customers read a third of the book in 2…

Cited by 15SourcePDFScholar
2018

LeapsAndBounds: A Method for Approximately Optimal Algorithm Configuration

ICML 2018oral

We consider the problem of configuring general-purpose solvers to run efficiently on problem instances drawn from an unknown distribution. The goal of the configurator is to find a configuration that runs fast on average on most instances, and do so with the least amount of total work. It can run a…

Cited by 46SourcePDFScholar
2015

On Identifying Good Options under Combinatorially Structured Feedback in Finite Noisy Environments

ICML 2015poster

We consider the problem of identifying a good option out of finite set of options under combinatorially structured, noisy feedback about the quality of the options in a sequential process: In each round, a subset of the options, from an available set of subsets, can be selected to receive noisy info…

Cited by 12SourcePDFScholar