2026
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future
ICML 2026poster
Self-Rewarding Language Models propose an architecture in which the Large Language Models(LLMs) both generates responses and evaluates its own outputs via LLM-as-a-Judge prompting, dynamically improving its generative capabilities through iterative Direct Preference Optimization (DPO). However, our …