← Search

Fabian Schramm

6 accepted papers

2026

Guided Flow Policy: Learning from High-Value Actions in Offline Reinforcement Learning

ICLR 2026poster

Offline reinforcement learning often relies on behavior regularization that enforces policies to remain close to the dataset distribution. However, such approaches fail to distinguish between high-value and low-value actions in their regularization components. We introduce Guided Flow Policy (GFP),…

Cited by 5SourcecodeScholar
2026

Reference-Free Sampling-Based Model Predictive Control

ICRA 2026poster

We present a sampling-based model predictive control (MPC) framework that enables emergent locomotion without relying on handcrafted gait patterns or predefined contact sequences. Our method discovers diverse motion patterns, ranging from trotting to galloping, robust standing policies, jumping, and…

2026

SVL: Goal-Conditioned Reinforcement Learning as Survival Learning

ICML 2026poster

Standard approaches to goal-conditioned reinforcement learning (GCRL) that rely on temporal-difference learning can be unstable and sample-inefficient due to bootstrapping. While recent work has explored contrastive and supervised formulations to improve stability, we present a probabilistic alterna…

Cited by 0SourceScholar
2026

Variance-Reduced Model Predictive Path Integral via Quadratic Model Approximation

RSS 2026poster

Sampling-based controllers, such as Model Predictive Path Integral (MPPI) methods, offer substantial flexibility but often suffer from high variance and low sample efficiency. To address these challenges, we introduce a hybrid variance-reduced MPPI framework that integrates a prior model into the sa…

Cited by 0SourceScholar
2024

Leveraging augmented-Lagrangian techniques for differentiating over infeasible quadratic programs in machine learning

ICLR 2024spotlight

Optimization layers within neural network architectures have become increasingly popular for their ability to solve a wide range of machine learning tasks and to model domain-specific knowledge. However, designing optimization layers requires careful consideration as the underlying optimization prob…

Cited by 3SourcePDFScholar
2022

Reactive Stepping for Humanoid Robots using Reinforcement Learning: Application to Standing Push Recovery on the Exoskeleton Atalante

IROS 2022poster

State-of-the-art reinforcement learning is now able to learn versatile locomotion, balancing and push-recovery capabilities for bipedal robots in simulation. Yet, the reality gap has mostly been overlooked and the simulated results hardly transfer to real hardware. Either it is unsuccessful in pract…

Cited by 13SourceScholar