← Search

Ronald Parr

6 accepted papers

2025

A Unifying View of Linear Function Approximation in Off-Policy RL Through Matrix Splitting and Preconditioning

NeurIPS 2025spotlight

In off-policy policy evaluation (OPE) tasks within reinforcement learning, Temporal Difference Learning(TD) and Fitted Q-Iteration (FQI) have traditionally been viewed as differing in the number of updates toward the target value function: TD makes one update, FQI makes an infinite number, and Parti…

Cited by 0SourceScholar
2024

Mitigating Partial Observability in Sequential Decision Processes via the Lambda Discrepancy

NeurIPS 2024poster

Reinforcement learning algorithms typically rely on the assumption that the environment dynamics and value function can be expressed in terms of a Markovian state representation. However, when state information is only partially observable, how can an agent learn such a state representation, and how…

2024

Position: Amazing Things Come From Having Many Good Models

ICML 2024spotlight

The *Rashomon Effect*, coined by Leo Breiman, describes the phenomenon that there exist many equally good predictive models for the same dataset. This phenomenon happens for many real datasets and when it does, it sparks both magic and consternation, but mostly magic. In light of the Rashomon Effect…

Cited by 25SourcePDFScholar
2024

Using Noise to Infer Aspects of Simplicity Without Learning

NeurIPS 2024poster

Noise in data significantly influences decision-making in the data science process. In fact, it has been shown that noise in data generation processes leads practitioners to find simpler models. However, an open question still remains: what is the degree of model simplification we can expect under d…

Cited by 0SourcePDFScholar