2026
Unified Generation and Self-Verification for Vision-Language Models via Advantage Decoupled Preference Optimization
CVPR 2026
Parallel test-time scaling typically trains separate generation and verification models, incurring high training and inference costs. We propose Advantage Decoupled Preference Optimization (ADPO), a unified reinforcement learning framework that jointly learns answer generation and self-verification