← Search

Sushant Prakash

3 accepted papers

2025

RLPF: Reinforcement Learning from Prediction Feedback for User Summarization with LLMs

AAAI 2025technical

LLM-powered personalization agent systems employ Large Language Models (LLMs) to predict users’ behavior from their past activities. However, their effectiveness often hinges on the ability to effectively leverage extensive, long user historical data due to its inherent noise and length of such data…

Cited by 2SourcePDFScholar
2024

RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

ICML 2024poster

Reinforcement learning from human feedback (RLHF) has proven effective in aligning large language models (LLMs) with human preferences, but gathering high-quality preference labels is expensive. RL from AI Feedback (RLAIF), introduced in Bai et al. (2022b), offers a promising alternative that trains…

Cited by 99SourcePDFScholar
2021

Federated Reconstruction: Partially Local Federated Learning

NeurIPS 2021poster

Personalization methods in federated learning aim to balance the benefits of federated and local training for data availability, communication cost, and robustness to client heterogeneity. Approaches that require clients to communicate all model parameters can be undesirable due to privacy and commu…