← Search

Pan Lihu

1 accepted papers

2026

Graph-Preference Learning: Debiasing Network-Sampled Human Feedback for Target Welfare Estimation

ICML 2026poster

Preference-based reward modeling is a core component of RLHF and DPO pipelines. In practice, the humans providing preference feedback are rarely an i.i.d. sample: recruitment and exposure often follow social, institutional, or spatial structure, inducing non-uniform inclusion probabilities that corr…

Cited by 0SourceScholar