← Search

Nikos Vlassis

9 accepted papers

2026

Stepwise Credit Assignment for GRPO on Flow-Matching Models

CVPR 2026

Flow-GRPO successfully applies reinforcement learning to flow models, but uses uniform credit assignment across all steps. This ignores the temporal structure of diffusion generation: early steps determine composition and content (low-frequency structure), while late steps resolve details and textur

Cited by 0SourceScholar
2024

Distributional Off-Policy Evaluation for Slate Recommendations

AAAI 2024technical

Recommendation strategies are typically evaluated by using previously logged data, employing off-policy evaluation methods to estimate their expected performance. However, for strategies that present users with slates of multiple items, the resulting combinatorial action space renders many of these…

2021

Control Variates for Slate Off-Policy Evaluation

NeurIPS 2021poster

We study the problem of off-policy evaluation from batched contextual bandit data with multidimensional actions, often termed slates. The problem is common to recommender systems and user-interface optimization, and it is particularly challenging because of the combinatorially-sized action space. Sw…

2019

More Efficient Off-Policy Evaluation through Regularized Targeted Learning

ICML 2019oral

We study the problem of off-policy evaluation (OPE) in Reinforcement Learning (RL), where the aim is to estimate the performance of a new policy given historical data that may have been generated by a different policy, or policies. In particular, we introduce a novel doubly-robust estimator for the…

Cited by 42SourcePDFScholar
2019

On the Design of Estimators for Bandit Off-Policy Evaluation

ICML 2019oral

Off-policy evaluation is the problem of estimating the value of a target policy using data collected under a different policy. Given a base estimator for bandit off-policy evaluation and a parametrized class of control variates, we address the problem of computing a control variate in that class tha…

Cited by 36SourcePDFScholar
2019

Optimizing over a Restricted Policy Class in MDPs

AISTATS 2019poster

We address the problem of finding an optimal policy in a Markov decision process (MDP) under a restricted policy class defined by the convex hull of a set of base policies. This problem is of great interest in applications in which a number of reasonably good (or safe) policies are already known and…

Cited by 9SourcePDFScholar