← Search

Riccardo De Santi

14 accepted papers

2026

A Unified Density Operator View of Flow Control and Merging

ICML 2026poster

Recent progress in large-scale flow and diffusion models raised two fundamental algorithmic challenges: $(i)$ control-based reward adaptation of pre-trained flows, and $(ii)$ integration of multiple models, i.e., flow merging. While current approaches address them separately, we introduce a unifying…

Cited by 0SourceScholar
2026

Constrained Flow Optimization via Sequential Fine-Tuning for Molecular Design

ICML 2026poster

Adapting generative foundation models, in particular diffusion and flow models, to optimize given reward functions (e.g., binding affinity) while satisfying constraints (e.g., molecular synthesizability) is fundamental for their adoption in real-world scientific discovery applications such as molecu…

Cited by 0SourceScholar
2026

Efficient Tail-Aware Generative Optimization via Flow Model Fine-Tuning

ICML 2026poster

Fine-tuning pre-trained diffusion and flow models to optimize downstream utilities is central to real-world deployment. Existing entropy-regularized methods primarily maximize expected reward, providing no mechanism to shape tail behavior. However, tail control is often essential: the lower tail det…

Cited by 0SourceScholar
2026

Flow Expansion via Verifier-Constrained Noised State Space Exploration

ICLR 2026poster

Flow and diffusion models are typically pre-trained on limited available data (e.g., molecular samples), covering only a fraction of the valid design space (e.g., the full molecular space). As a consequence, they tend to generate samples from only a narrow portion of the feasible domain. This is a f…

Cited by 0SourceScholar
2026

Landing with the Score: Riemannian Optimization through Denoising

ICLR 2026poster

Under the \emph{data manifold hypothesis}, high-dimensional data concentrate near a low-dimensional manifold. We study Riemannian optimization when this manifold is only given implicitly through the data distribution, and standard geometric operations are unavailable. This formulation captures a br…

Cited by 0SourceScholar
2026

Value Matching: Scalable and Gradient-Free Reward-Guided Flow Adaptation

ICLR 2026poster

Adapting large-scale flow and diffusion models to downstream tasks through reward optimization is essential for their adoption in real-world applications, including scientific discovery and image generation. While recent fine-tuning methods based on reinforcement learning and stochastic optimal cont…

Cited by 0SourceScholar
2025

Flow Density Control: Generative Optimization Beyond Entropy-Regularized Fine-Tuning

NeurIPS 2025spotlight

Adapting large-scale foundational flow and diffusion generative models to optimize task-specific objectives while preserving prior information is crucial for real-world applications such as molecular design, protein docking, and creative image generation. Existing principled fine-tuning methods aim…

Cited by 0SourceScholar
2025

Provable Maximum Entropy Manifold Exploration via Diffusion Models

ICML 2025poster

Exploration is critical for solving real-world decision-making problems such as scientific discovery, where the objective is to generate truly novel designs rather than mimic existing data distributions. In this work, we address the challenge of leveraging the representational power of generative m…

Cited by 0SourcePDFScholar
2024

Exploiting Causal Graph Priors with Posterior Sampling for Reinforcement Learning

ICLR 2024poster

Posterior sampling allows exploitation of prior knowledge on the environment's transition dynamics to improve the sample efficiency of reinforcement learning. The prior is typically specified as a class of parametric distributions, the design of which can be cumbersome in practice, often resulting i…

Cited by 5SourcePDFScholar
2024

Geometric Active Exploration in Markov Decision Processes: the Benefit of Abstraction

ICML 2024poster

How can a scientist use a Reinforcement Learning (RL) algorithm to design experiments over a dynamical system's state space? In the case of finite and Markovian systems, an area called *Active Exploration* (AE) relaxes the optimization problem of experiments design into Convex RL, a generalization o…

Cited by 1SourcePDFScholar
2024

Global Reinforcement Learning : Beyond Linear and Convex Rewards via Submodular Semi-gradient Methods

ICML 2024poster

In classic Reinforcement Learning (RL), the agent maximizes an additive objective of the visited states, e.g., a value function. Unfortunately, objectives of this type cannot model many real-world applications such as experiment design, exploration, imitation learning, and risk-averse RL to name a f…

Cited by 7SourcePDFScholar
2023

Provably Efficient Causal Model-Based Reinforcement Learning for Systematic Generalization

AAAI 2023technical

In the sequential decision making setting, an agent aims to achieve systematic generalization over a large, possibly infinite, set of environments. Such environments are modeled as discrete Markov decision processes with both states and actions represented through a feature vector. The underlying st…

Cited by 19SourcePDFScholar
2022

Challenging Common Assumptions in Convex Reinforcement Learning

NeurIPS 2022accept

The classic Reinforcement Learning (RL) formulation concerns the maximization of a scalar reward function. More recently, convex RL has been introduced to extend the RL formulation to all the objectives that are convex functions of the state distribution induced by a policy. Notably, convex RL cover…

Cited by 27SourcePDFScholar
2022

The Importance of Non-Markovianity in Maximum State Entropy Exploration

ICML 2022oral

In the maximum state entropy exploration framework, an agent interacts with a reward-free environment to learn a policy that maximizes the entropy of the expected state visitations it is inducing. Hazan et al. (2019) noted that the class of Markovian stochastic policies is sufficient for the maximum…

Cited by 41SourcePDFScholar