← Search

Florian Dorfler

12 accepted papers

2026

Landing with the Score: Riemannian Optimization through Denoising

ICLR 2026poster

Under the \emph{data manifold hypothesis}, high-dimensional data concentrate near a low-dimensional manifold. We study Riemannian optimization when this manifold is only given implicitly through the data distribution, and standard geometric operations are unavailable. This formulation captures a br…

Cited by 0SourceScholar
2026

Sample-efficient and Scalable Exploration in Continuous-Time RL

ICLR 2026poster

Reinforcement learning algorithms are typically designed for discrete-time dynamics, even though the underlying real-world control systems are often continuous in time. In this paper, we study the problem of continuous-time reinforcement learning, where the unknown system dynamics are represented us…

Cited by 0SourcecodeScholar
2025

Contractivity and linear convergence in bilinear saddle-point problems: An operator-theoretic approach

AISTATS 2025poster

We study the convex-concave bilinear saddle-point problem $\min_x \max_y f(x) + y^\top Ax - g(y)$, where both, only one, or none of the functions $f$ and $g$ are strongly convex, and suitable rank conditions on the matrix $A$ hold. The solution of this problem is at the core of many machine learning…

Cited by 0SourceScholar
2025

Optimizing Social Network Interventions via Hypergradient-Based Recommender System Design

ICML 2025poster

Although social networks have expanded the range of ideas and information accessible to users, they are also criticized for amplifying the polarization of user opinions. Given the inherent complexity of these phenomena, existing approaches to counteract these effects typically rely on handcrafted al…

Cited by 0SourcePDFScholar
2025

SOMBRL: Scalable and Optimistic Model-Based RL

NeurIPS 2025poster

We address the challenge of efficient exploration in model-based reinforcement learning (MBRL), where the system dynamics are unknown and the RL agent must learn directly from online interactions. We propose **S**calable and **O**ptimistic **MBRL** (SOMBRL), an approach based on the principle of opt…

Cited by 0SourceScholar
2024

Fairness in Social Influence Maximization via Optimal Transport

NeurIPS 2024poster

We study fairness in social influence maximization, whereby one seeks to select seeds that spread a given information throughout a network, ensuring balanced outreach among different communities (e.g. demographic groups). In the literature, fairness is often quantified in terms of the expected outre…

2024

NeoRL: Efficient Exploration for Nonepisodic RL

NeurIPS 2024spotlight

We study the problem of nonepisodic reinforcement learning (RL) for nonlinear dynamical systems, where the system dynamics are unknown and the RL agent has to learn from a single trajectory, i.e., without resets. We propose **N**on**e**pisodic **O**ptistmic **RL** (NeoRL), an approach based on the p…

Cited by 2SourcePDFScholar
2024

When to Sense and Control? A Time-adaptive Approach for Continuous-Time RL

NeurIPS 2024poster

Reinforcement learning (RL) excels in optimizing policies for discrete-time Markov decision processes (MDP). However, various systems are inherently continuous in time, making discrete-time MDPs an inexact modeling choice. In many applications, such as greenhouse control or medical treatments, each…

2023

Efficient Exploration in Continuous-time Model-based Reinforcement Learning

NeurIPS 2023poster

Reinforcement learning algorithms typically consider discrete-time dynamics, even though the underlying systems are often continuous in time. In this paper, we introduce a model-based reinforcement learning algorithm that represents continuous-time dynamics using nonlinear ordinary differential equa…

Cited by 10SourcePDFScholar
2022

Factorization of Dynamic Games over Spatio-Temporal Resources

IROS 2022poster

Dynamic games feature a state-space complexity that scales superlinearly with the number of players. This makes this class of games often intractable even for a handful of players. We introduce the factorization process of dynamic games as a transformation leveraging the independence of players at e…

Cited by 7SourceScholar
2022

Trust Region Policy Optimization with Optimal Transport Discrepancies: Duality and Algorithm for Continuous Actions

NeurIPS 2022accept

Policy Optimization (PO) algorithms have been proven particularly suited to handle the high-dimensionality of real-world continuous control tasks. In this context, Trust Region Policy Optimization methods represent a popular approach to stabilize the policy updates. These usually rely on the Kullbac…

Cited by 13SourcePDFScholar