← Search

Marin Vlastelica

15 accepted papers

2026

Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead

ICLR 2026poster

In post-training for reasoning Large Language Models (LLMs), the current state of practice trains LLMs in two independent stages: Supervised Fine-Tuning (SFT) and Reinforcement Learning with Verifiable Rewards (RLVR, shortened as "RL" below). In this work, we challenge whether high SFT scores transl…

Cited by 0SourceScholar
2025

Divide, Discover, Deploy: Factorized Skill Learning with Symmetry and Style Priors

CoRL 2025oral

Unsupervised Skill Discovery (USD) allows agents to autonomously learn diverse behaviors without task-specific rewards. While recent USD methods have shown promise, their application to real-world robotics remains underexplored. In this paper, we propose a modular USD framework to address the challe…

Cited by 0SourceScholar
2025

Flow Density Control: Generative Optimization Beyond Entropy-Regularized Fine-Tuning

NeurIPS 2025spotlight

Adapting large-scale foundational flow and diffusion generative models to optimize task-specific objectives while preserving prior information is crucial for real-world applications such as molecular design, protein docking, and creative image generation. Existing principled fine-tuning methods aim…

Cited by 0SourceScholar
2025

Provable Maximum Entropy Manifold Exploration via Diffusion Models

ICML 2025poster

Exploration is critical for solving real-world decision-making problems such as scientific discovery, where the objective is to generate truly novel designs rather than mimic existing data distributions. In this work, we address the challenge of leveraging the representational power of generative m…

Cited by 0SourcePDFScholar
2024

Causal Action Influence Aware Counterfactual Data Augmentation

ICML 2024poster

Offline data are both valuable and practical resources for teaching robots complex behaviors. Ideally, learning agents should not be constrained by the scarcity of available demonstrations, but rather generalize beyond the training distribution. However, the complexity of real-world scenarios typica…

2024

Learning Diverse Skills for Local Navigation under Multi-constraint Optimality

ICRA 2024poster

Despite many successful applications of data-driven control in robotics, extracting meaningful diverse behaviors remains a challenge. Typically, task performance needs to be compromised in order to achieve diversity. In many scenarios, task requirements are specified as a multitude of reward terms,…

Cited by 7SourceScholar
2023

Backpropagation through Combinatorial Algorithms: Identity with Projection Works

ICLR 2023poster

Embedding discrete solvers as differentiable layers has given modern deep learning architectures combinatorial expressivity and discrete reasoning capabilities. The derivative of these solvers is zero or undefined, therefore a meaningful replacement is crucial for effective gradient-based learning.…

2023

Spuriosity Didn’t Kill the Classifier: Using Invariant Predictions to Harness Spurious Features

NeurIPS 2023poster

To avoid failures on out-of-distribution data, recent works have sought to extract features that have an invariant or stable relationship with the label across domains, discarding "spurious" or unstable features whose relationship with the label changes across domains. However, unstable features oft…

Cited by 19SourcePDFScholar
2023

Taming Continuous Posteriors for Latent Variational Dialogue Policies

AAAI 2023technical

Utilizing amortized variational inference for latent-action reinforcement learning (RL) has been shown to be an effective approach in Task-oriented Dialogue (ToD) systems for optimizing dialogue success.Until now, categorical posteriors have been argued to be one of the main drivers of performance.…

Cited by 2SourcePDFScholar
2023

Versatile Skill Control via Self-supervised Adversarial Imitation of Unlabeled Mixed Motions

ICRA 2023poster

Learning diverse skills is one of the main challenges in robotics. To this end, imitation learning approaches have achieved impressive results. These methods require explicitly labeled datasets or assume consistent skill execution to enable learning and active control of individual behaviors, which…

Cited by 34SourceScholar
2022

Learning Agile Skills via Adversarial Imitation of Rough Partial Demonstrations

CoRL 2022oral

Learning agile skills is one of the main challenges in robotics. To this end, reinforcement learning approaches have achieved impressive results. These methods require explicit task information in terms of a reward function or an expert that can be queried in simulation to provide a target control o…

Cited by 73SourceScholar
2021

Neuro-algorithmic Policies Enable Fast Combinatorial Generalization

ICML 2021spotlight

Although model-based and model-free approaches to learning the control of systems have achieved impressive results on standard benchmarks, generalization to task variations is still lacking. Recent results suggest that generalization for standard architectures improves only after obtaining exhaustiv…

Cited by 19SourcePDFScholar
2020

Optimizing Rank-Based Metrics With Blackbox Differentiation

CVPR 2020oral

Rank-based metrics are some of the most widely used criteria for performance evaluation of computer vision models. Despite years of effort, direct optimization for these metrics remains a challenge due to their non-differentiable and non-decomposable nature. We present an efficient, theoretically so…

Cited by 121PDFcodeScholar