2026
QUATRO: Query-Adaptive Trust Region Policy Optimization for LLM Fine-tuning
ICML 2026poster
GRPO-style reinforcement learning (RL)-based LLM fine-tuning algorithms have recently gained popularity. Relying on heuristic trust-region approximations, however, they can lead to brittle optimization behavior, as global importance-ratio clipping and group-wise normalization fail to regulate sample…