← Search

Daniel Roy

7 accepted papers

2021

On the Role of Data in PAC-Bayes Bounds

AISTATS 2021poster

The dominant term in PAC-Bayes bounds is often the Kullback-Leibler divergence between the posterior and prior. For so-called linear PAC-Bayes risk bounds based on the empirical risk of a fixed posterior kernel, it is possible to minimize the expected value of the bound by choosing the prior to be t…

2021

Pruning Neural Networks at Initialization: Why Are We Missing the Mark?

ICLR 2021poster

Recent work has explored the possibility of pruning neural networks at initialization. We assess proposals for doing so: SNIP (Lee et al., 2019), GraSP (Wang et al., 2020), SynFlow (Tanaka et al., 2020), and magnitude pruning. Although these methods surpass the trivial baseline of random pruning, th…

Cited by 275SourcePDFScholar
2020

In Defense of Uniform Convergence: Generalization via Derandomization with an Application to Interpolating Predictors

ICML 2020accepted

We propose to study the generalization error of a learned predictor in terms of that of a surrogate (potentially randomized) predictor that is coupled to $\hh$ and designed to trade empirical risk for control of generalization error. In the case where the learned predictor interpolates the data, it…

Cited by 74SourcePDFScholar
2020

Linear Mode Connectivity and the Lottery Ticket Hypothesis

ICML 2020poster

We study whether a neural network optimizes to the same, linearly connected minimum under different samples of SGD noise (e.g., random data order and augmentation). We find that standard vision models become stable to SGD noise in this way early in training. From then on, the outcome of optimization…

2020

Tight Bounds on Minimax Regret under Logarithmic Loss via Self-Concordance

ICML 2020poster

We consider the classical problem of sequential probability assignment under logarithmic loss while competing against an arbitrary, potentially nonparametric class of experts. We obtain tight bounds on the minimax regret via a new approach that exploits the self-concordance property of the logarithm…

Cited by 25SourcePDFScholar
2018

Entropy-SGD optimizes the prior of a PAC-Bayes bound: Generalization properties of Entropy-SGD and data-dependent priors

ICML 2018oral

We show that Entropy-SGD (Chaudhari et al., 2017), when viewed as a learning algorithm, optimizes a PAC-Bayes bound on the risk of a Gibbs (posterior) classifier, i.e., a randomized classifier obtained by a risk-sensitive perturbation of the weights of a learned classifier. Entropy-SGD works by opti…

Cited by 66SourcePDFScholar