2025
Generative RLHF-V: Learning Principles from Multi-modal Human Preference
NeurIPS 2025poster
Training multi-modal large language models (MLLMs) that align with human intentions is a long-term challenge. Traditional score-only reward models for alignment suffer from low accuracy, weak generalization, and poor interpretability, blocking the progress of alignment methods, \textit{e.g.,} reinfo…