2026
Translate Policy to Language: Flow Matching Generated Rewards for LLM Explanations
ICLR 2026poster
As humans increasingly share environments with diverse agents powered by RL, LLMs, and beyond, the ability to explain agent policies in natural language is vital for reliable coexistence. We introduce a general-purpose framework that trains explanation-generating LLMs via reinforcement learning from…