← Search

Bibek Aryal

1 accepted papers

2025

RLTHF: Targeted Human Feedback for LLM Alignment

ICML 2025poster

Fine-tuning large language models (LLMs) to align with user preferences is challenging due to the high cost of quality human annotations in Reinforcement Learning from Human Feedback (RLHF) and the generalizability limitations of AI Feedback. To address these challenges, we propose RLTHF, a human-AI…

Cited by 0SourcePDFScholar