DiffOP: Reinforcement Learning of Optimization-Based Control Policies via Implicit Policy Gradients
Real-world control systems require policies that are not only high-performing but also interpretable and robust. A promising direction toward this goal is model-based control, which learns system dynamics and cost functions from historical data and then uses these models to inform decision-making. B