2025
Reward Mixology: Crafting Hybrid Signals for Reinforcement Learning Driven In-Context Learning
EMNLP 2025
In-context learning (ICL) performance heavily relies on the quality and ordering of demonstrations. Iterative selection (IS) is a promising approach to address this issue, but existing IS methods face two key challenges: the oversimplification of process reward signals that guide intermediate steps