← Search

Chu Zhao

2 accepted papers

2026

Causal Direct Preference Optimization for Distributionally Robust Generative Recommendation

ICML 2026poster

Direct Preference Optimization (DPO) guides large language models (LLMs) to generate recommendations aligned with user historical behavior distributions by minimizing preference alignment loss. However, our systematic empirical research and theoretical analysis reveal that DPO tends to amplify spuri…

Cited by 0SourceScholar
2026

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning

ICML 2026poster

Test-time reinforcement learning generates multiple candidate answers via repeated rollouts and performs online updates using pseudo-labels constructed by majority voting. To reduce overhead and improve exploration, prior work introduces tree-structured rollouts, which share reasoning prefixes and b…

Cited by 0SourceScholar