← Search

Luckeciano Carvalho Melo

4 accepted papers

2026

Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning

ICLR 2026poster

Reinforcement Learning, particularly through policy gradient methods, has played a central role in enabling reasoning capabilities of Large Language Models. However, the optimization stability of policy gradients in this setting remains understudied. As a result, existing implementations often resor…

Cited by 0SourcecodeScholar
2025

Sliding Puzzles Gym: A Scalable Benchmark for State Representation in Visual Reinforcement Learning

ICML 2025poster

Effective visual representation learning is crucial for reinforcement learning (RL) agents to extract task-relevant information from raw sensory inputs and generalize across diverse environments. However, existing RL benchmarks lack the ability to systematically evaluate representation learning capa…

2024

Deep Bayesian Active Learning for Preference Modeling in Large Language Models

NeurIPS 2024poster

Leveraging human preferences for steering the behavior of Large Language Models (LLMs) has demonstrated notable success in recent years. Nonetheless, data selection and labeling are still a bottleneck for these systems, particularly at large scale. Hence, selecting the most informative points for ac…