LeapAlign: Post-training Flow Matching Models at Any Generation Step by Building Two-Step Trajectories
This paper focuses on the alignment of flow-matching models with human preference. A promising way is fine-tuning by directly backpropagating reward signals through the differentiable generation process of flow matching. However, backpropagating through long trajectories results in prohibitive memor