← Search

Viraj Mehta

8 accepted papers

2024

Group Robust Preference Optimization in Reward-free RLHF

NeurIPS 2024poster

Adapting large language models (LLMs) for specific tasks usually involves fine-tuning through reinforcement learning with human feedback (RLHF) on preference data. While these data often come from diverse labelers' groups (e.g., different demographics, ethnicities, company teams, etc.), traditional…

Cited by 20SourcePDFScholar
2024

Position: Opportunities Exist for Machine Learning in Magnetic Fusion Energy

ICML 2024oral

Magnetic confinement fusion may one day provide reliable, carbon-free energy, but the field currently faces technical hurdles. In this position paper, we highlight six key research challenges in the field of fusion energy that we believe should be research priorities for the Machine Learning (ML) co…

Cited by 1SourcePDFScholar
2023

Near-optimal Policy Identification in Active Reinforcement Learning

ICLR 2023top-5%

Many real-world reinforcement learning tasks require control of complex dynamical systems that involve both costly data acquisition processes and large state spaces. In cases where the expensive transition dynamics can be readily evaluated at specified states (e.g., via a simulator), agents can oper…

Cited by 8SourcePDFScholar
2022

An Experimental Design Perspective on Model-Based Reinforcement Learning

ICLR 2022poster

In many practical applications of RL, it is expensive to observe state transitions from the environment. For example, in the problem of plasma control for nuclear fusion, computing the next state for a given state-action pair requires querying an expensive transition function which can lead to many…

Cited by 36SourcePDFScholar
2022

Exploration via Planning for Information about the Optimal Trajectory

NeurIPS 2022accept

Many potential applications of reinforcement learning (RL) are stymied by the large numbers of samples required to learn an effective policy. This is especially true when applying RL to real-world control tasks, e.g. in the sciences or robotics, where executing a policy in the environment is costly.…

2022

Variational autoencoders in the presence of low-dimensional data: landscape and implicit bias

ICLR 2022poster

Variational Autoencoders (VAEs) are one of the most commonly used generative models, particularly for image data. A prominent difficulty in training VAEs is data that is supported on a lower dimensional manifold. Recent work by Dai and Wipf (2020) proposes a two-stage training algorithm for VAEs, ba…

2021

Representational aspects of depth and conditioning in normalizing flows

ICML 2021spotlight

Normalizing flows are among the most popular paradigms in generative modeling, especially for images, primarily because we can efficiently evaluate the likelihood of a data point. This is desirable both for evaluating the fit of a model, and for ease of training, as maximizing the likelihood can be…

Cited by 40SourcePDFScholar
2018

Learning Task-Oriented Grasping for Tool Manipulation from Simulated Self-Supervision

RSS 2018poster

Tool manipulation is vital for facilitating robots to complete challenging task goals. It requires reasoning about the desired effect of the task and thus properly grasping and manipulating the tool to achieve the task. Task-agnostic grasping optimizes for grasp robustness while ignoring crucial tas…

Cited by 259SourcePDFScholar