← Search

Ryan J Tibshirani

9 accepted papers

2019

A Continuous-Time View of Early Stopping for Least Squares Regression

AISTATS 2019poster

We study the statistical properties of the iterates generated by gradient descent, applied to the fundamental problem of least squares regression. We take a continuous-time view, i.e., consider infinitesimal step sizes in gradient descent, in which case the iterates form a trajectory called gradient…

Cited by 157SourcePDFScholar
2019

A Higher-Order Kolmogorov-Smirnov Test

AISTATS 2019poster

We present an extension of the Kolmogorov-Smirnov (KS) two-sample test, which can be more sensitive to differences in the tails. Our test statistic is an integral probability metric (IPM) defined over a higher-order total variation ball, recovering the original KS test as its simplest case. We giv…

Cited by 18SourcePDFScholar
2019

Conformal Prediction Under Covariate Shift

NeurIPS 2019poster

We extend conformal prediction methodology beyond the case of exchangeable data. In particular, we show that a weighted version of conformal prediction can be used to compute distribution-free prediction intervals for problems in which the test and training covariate distributions differ, but the li…

2019

Kalman Filter, Sensor Fusion, and Constrained Regression: Equivalences and Insights

NeurIPS 2019poster

The Kalman filter (KF) is one of the most widely used tools for data assimilation and sequential estimation. In this work, we show that the state estimates from the KF in a standard linear dynamical system setting are equivalent to those given by the KF in a transformed system, with infinite process…

2017

A Sharp Error Analysis for the Fused Lasso, with Application to Approximate Changepoint Screening

NeurIPS 2017poster

In the 1-dimensional multiple changepoint detection problem, we derive a new fast error rate for the fused lasso estimator, under the assumption that the mean vector has a sparse number of changepoints. This rate is seen to be suboptimal (compared to the minimax rate) by only a factor of $\log\log{n…

Cited by 77SourcePDFScholar
2017

Higher-Order Total Variation Classes on Grids: Minimax Theory and Trend Filtering Methods

NeurIPS 2017poster

We consider the problem of estimating the values of a function over $n$ nodes of a $d$-dimensional grid graph (having equal side lengths $n^{1/d}$) from noisy observations. The function is assumed to be smooth, but is allowed to exhibit different amounts of smoothness at different regions in the gri…

Cited by 39SourcePDFScholar
2016

Total Variation Classes Beyond 1d: Minimax Rates, and the Limitations of Linear Smoothers

NeurIPS 2016poster

We consider the problem of estimating a function defined over $n$ locations on a $d$-dimensional grid (having all side lengths equal to $n^{1/d}$). When the function is constrained to have discrete total variation bounded by $C_n$, we derive the minimax optimal (squared) $\ell_2$ estimation error r…

Cited by 92SourcePDFScholar