← Search

Ujwal Dinesha

2 accepted papers

2025

DOPL: Direct Online Preference Learning for Restless Bandits with Preference Feedback

ICLR 2025poster

Restless multi-armed bandits (RMAB) has been widely used to model constrained sequential decision making problems, where the state of each restless arm evolves according to a Markov chain and each state transition generates a scalar reward. However, the success of RMAB crucially relies on the availa…

Cited by 2SourcePDFScholar
2024

Risk-Averse Fine-tuning of Large Language Models

NeurIPS 2024poster

We consider the challenge of mitigating the generation of negative or toxic content by the Large Language Models (LLMs) in response to certain prompts. We propose integrating risk-averse principles into LLM fine-tuning to minimize the occurrence of harmful outputs, particularly rare but significant…