← Search

Xiaoran Shi

1 accepted papers

2025

GAPO: Learning Preferential Prompt through Generative Adversarial Policy Optimization

ACL 2025long

Recent advances in large language models have highlighted the critical need for precise control over model outputs through predefined constraints. While existing methods attempt to achieve this through either direct instruction-response synthesis or preferential response optimization, they often str…