← Search

Daniel Melcer

3 accepted papers

2025

Approximately Aligned Decoding

NeurIPS 2025poster

It is common to reject undesired outputs of Large Language Models (LLMs); however, current methods to do so require an excessive amount of computation to re-sample after a rejection, or distort the distribution of outputs by constraining the output to highly improbable tokens. We present a method, A…

Cited by 0SourceScholar
2022

Shield Decentralization for Safe Multi-Agent Reinforcement Learning

NeurIPS 2022accept

Learning safe solutions is an important but challenging problem in multi-agent reinforcement learning (MARL). Shielded reinforcement learning is one approach for preventing agents from choosing unsafe actions. Current shielded reinforcement learning methods for MARL make strong assumptions about com…

Cited by 20SourcePDFScholar
2021

Dynamic Automaton-Guided Reward Shaping for Monte Carlo Tree Search

AAAI 2021technical

Reinforcement learning and planning have been revolutionized in recent years, due in part to the mass adoption of deep convolutional neural networks and the resurgence of powerful methods to refine decision-making policies. However, the problem of sparse reward signals and their representation remai…

Cited by 22SourcePDFScholar