← Search

Matthew Cleaveland

5 accepted papers

2026

ReFORM: Reflected Flows for On-support Offline RL via Noise Manipulation

ICLR 2026poster

Offline reinforcement learning (RL) aims to learn the optimal policy from a fixed dataset generated by behavior policies without additional environment interactions. One common challenge that arises in this setting is the out-of-distribution (OOD) error, which occurs when the policy leaves the train…

Cited by 0SourcecodeScholar
2026

Solving Parameter-Robust Avoid Problems with Unknown Feasibility using Reinforcement Learning

ICLR 2026poster

Recent advances in deep reinforcement learning (RL) have achieved strong results on high-dimensional control tasks, but applying RL to reachability problems raises a fundamental mismatch: reachability seeks to maximize the set of states from which a system remains safe indefinitely, while RL optimiz…

Cited by 0SourceScholar
2024

Conformal Prediction Regions for Time Series Using Linear Complementarity Programming

AAAI 2024technical

Conformal prediction is a statistical tool for producing prediction regions of machine learning models that are valid with high probability. However, applying conformal prediction to time series data leads to conservative prediction regions. In fact, to obtain prediction regions over T time steps…

2023

Safe Planning in Dynamic Environments Using Conformal Prediction

RA-L 2023

We propose a framework for planning in unknown dynamic environments with probabilistic safety guarantees using conformal prediction. Particularly, we design a model predictive controller (MPC) that uses i) trajectory predictions of the dynamic environment, and ii) prediction regions quantifying the

Cited by 195SourceScholar
2022

Learning Enabled Fast Planning and Control in Dynamic Environments with Intermittent Information

IROS 2022poster

This paper addresses a safe planning and control problem for mobile robots operating in communication- and sensor-limited dynamic environments. In this case the robots cannot sense the objects around them and must instead rely on intermittent, external information about the environment, as e.g., in…

Cited by 1SourceScholar