GAPO: Learning Preferential Prompt through Generative Adversarial Policy Optimization
Recent advances in large language models have highlighted the critical need for precise control over model outputs through predefined constraints. While existing methods attempt to achieve this through either direct instruction-response synthesis or preferential response optimization, they often str…