← Search

Yanming Wan

5 accepted papers

2026

Learning to summarize user information for personalized reinforcement learning from human feedback

ICLR 2026poster

As everyday use cases of large language model (LLM) AI assistants have expanded, it is becoming increasingly important to personalize responses to align to different users' preferences and goals. While reinforcement learning from human feedback (RLHF) is effective at improving LLMs to be generally m…

Cited by 0SourceScholar
2025

Enhancing Personalized Multi-Turn Dialogue with Curiosity Reward

NeurIPS 2025poster

Effective conversational agents must personalize their interactions to adapt to user preferences, personalities, and attributes across diverse domains like education and healthcare. Current methods like Reinforcement Learning from Human Feedback (RLHF), often prioritize helpfulness and safety but fa…

Cited by 0SourceScholar
2025

Infer Human’s Intentions Before Following Natural Language Instructions

AAAI 2025technical

For AI agents to be helpful to humans, they should be able to follow natural language instructions to complete everyday cooperative tasks in human environments. However, real human instructions inherently possess ambiguity, because the human speakers assume sufficient prior knowledge about their hid…

2024

Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning

NeurIPS 2024spotlight

Reinforcement Learning from Human Feedback (RLHF) is a powerful paradigm for aligning foundation models to human values and preferences. However, current RLHF techniques cannot account for the naturally occurring differences in individual human preferences across a diverse population. When these dif…

Cited by 29SourcePDFScholar
2022

HandMeThat: Human-Robot Communication in Physical and Social Environments

NeurIPS 2022accept

We introduce HandMeThat, a benchmark for a holistic evaluation of instruction understanding and following in physical and social environments. While previous datasets primarily focused on language grounding and planning, HandMeThat considers the resolution of human instructions with ambiguities base…

Cited by 21SourcePDFScholar