2026
Learning from Comparison: Constrained Projection Policy Optimization for Pareto-Front Improvement
ICML 2026poster
Constrained multi-objective reinforcement learning aims to discover a diverse set of feasible trade-offs, yet scalarization and signed, normalized group-relative advantages can be brittle under objective-scale drift, near-ties, and feasibility scarcity. We propose constrained projection policy optim…