← Search

Daniel Marta

9 accepted papers

2026

Reinforcement Learning via Self-Distillation

ICML 2026poster

Large language models are increasingly post-trained with reinforcement learning in verifiable domains such as code and math. Yet, current methods for reinforcement learning with verifiable rewards (RLVR) learn only from a scalar outcome reward per attempt, creating a severe credit-assignment bottlen…

Cited by 0SourceScholar
2025

Flora: Sample-Efficient Preference-Based Rl Via Low-Rank Style Adaptation of Reward Functions

ICRA 2025

Preference-based reinforcement learning (PbRL) is a suitable approach for style adaptation of pre-trained robotic behavior: adapting the robot's policy to follow human user preferences while still being able to perform the original task. However, collecting preferences for the adaptation process in

Cited by 2SourcecodeScholar
2025

The Impact of VR and 2D Interfaces on Human Feedback in Preference-Based Robot Learning

IROS 2025

Aligning robot navigation with human preferences is essential for ensuring comfortable, and predictable robot movement in shared spaces. While preference-based learning methods, such as reinforcement learning from human feedback (RLHF), enable this alignment, the choice of the preference collection

Cited by 3SourceScholar
2024

SEQUEL: Semi-Supervised Preference-based RL with Query Synthesis via Latent Interpolation

ICRA 2024poster

Preference-based reinforcement learning (RL) poses as a recent research direction in robot learning, by allowing humans to teach robots through preferences on pairs of desired behaviours. Nonetheless, to obtain realistic robot policies, an arbitrarily large number of queries is required to be answer…

Cited by 3SourceScholar
2023

Aligning Human Preferences with Baseline Objectives in Reinforcement Learning

ICRA 2023poster

Practical implementations of deep reinforcement learning (deep RL) have been challenging due to an amplitude of factors, such as designing reward functions that cover every possible interaction. To address the heavy burden of robot reward engineering, we aim to leverage subjective human preferences…

Cited by 17SourceScholar
2023

VARIQuery: VAE Segment-Based Active Learning for Query Selection in Preference-Based Reinforcement Learning

IROS 2023poster

Human-in-the-loop reinforcement learning (RL) methods actively integrate human knowledge to create reward functions for various robotic tasks. Learning from preferences shows promise as alleviates the requirement of demonstrations by querying humans on state-action sequences. However, the limited gr…

Cited by 8SourceScholar
2022

Human-Feedback Shield Synthesis for Perceived Safety in Deep Reinforcement Learning

RA-L 2022

Despite the successes of deep reinforcement learning (RL), it is still challenging to obtain safe policies. Formal verification approaches ensure safety at all times, but usually overly restrict the agent’s behaviors, since they assume adversarial behavior of the environment. Instead of assuming adv

Cited by 14SourceScholar