← Search

Liran Ma

1 accepted papers

2026

Reliability-Aware LLM Alignment from Inconsistent Human Feedback

ICML 2026poster

Reinforcement Learning from Human Feedback (RLHF) is critical for aligning Large Language Models (LLMs) with human preferences. However, its efficacy is often compromised by the inherent inconsistency and subjectivity of human annotations. Existing preference optimization frameworks, such as Direct …

Cited by 0SourceScholar