← Search

Guofeng Cui

5 accepted papers

2026

Preference-based Policy Optimization from Sparse-reward Offline Dataset

ICLR 2026poster

Offline reinforcement learning (RL) holds the promise of training effective policies from static datasets without the need for costly online interactions. However, offline RL faces key limitations, most notably the challenge of generalizing to unseen or infrequently encountered state-action pairs. W…

Cited by 0SourceScholar
2025

CRPO: Confidence-Reward Driven Preference Optimization for Machine Translation

ACL 2025finding

Large language models (LLMs) have shown great potential in natural language processing tasks, but their application to machine translation (MT) remains challenging due to pretraining on English-centric data and the complexity of reinforcement learning from human feedback (RLHF). Direct Preference Op…

Cited by 0SourcePDFScholar
2024

Exploring the Edges of Latent State Clusters for Goal-Conditioned Reinforcement Learning

NeurIPS 2024poster

Exploring unknown environments efficiently is a fundamental challenge in unsupervised goal-conditioned reinforcement learning. While selecting exploratory goals at the frontier of previously explored states is an effective strategy, the policy during training may still have limited capability of rea…

2020

Talking-head Generation with Rhythmic Head Motion

ECCV 2020poster

When people deliver a speech, they naturally move heads, and this rhythmic head motion conveys linguistic information. However, generating a lip-synced video while moving head naturally is challenging. While remarkably successful, existing works either generate still talking-face videos or rely on l…