NeurIPS 2025poster0 citations

Retrospective In-Context Learning for Temporal Credit Assignment with Large Language Models

Wentse Chen, Jiayu Chen, Fahim Tajwar, Hao Zhu, Xintong Duan, Ruslan Salakhutdinov, Jeff Schneider

Abstract

Learning from self-sampled data and sparse environmental feedback remains a fundamental challenge in training self-evolving agents. Temporal credit assignment mitigates this issue by transforming sparse feedback into dense supervision signals. However, previous approaches typically depend on domain-specific value functions for credit assignment, which suffer from poor sample efficiency and limited generalization. In this work, we propose to leverage pre-trained knowledge from large language models (LLMs) to transform sparse rewards into dense training signals (i.e., the advantage function) through retrospective in-context learning (RICL). We further propose an online learning framework, RICOL, which iteratively refines the policy based on the credit assignment results from RICL. We empirically demonstrate that RICL can accurately estimate the advantage function with limited samples and effectively identify critical states for temporal credit assignment. Extended evaluation on the BabyAI benchmark shows that RICOL significantly improves sample efficiency compared to traditional online RL algorithms while achieving performance comparable to imitation learning from expert demonstartions. Our findings highlight the potential of leveraging LLMs for temporal credit assignment, paving the way for more sample-efficient and generalizable RL paradigms.

In-Context LearningCredit AssignmentLarge Language ModelsReinforcement Learning
BibTeX
@inproceedings{
chen2025retrospective,
title={Retrospective In-Context Learning for Temporal Credit Assignment with Large Language Models},
author={Wentse Chen and Jiayu Chen and Fahim Tajwar and Hao Zhu and Xintong Duan and Ruslan Salakhutdinov and Jeff Schneider},
booktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},
year={2025},
url={https://openreview.net/forum?id=QAVpe6a3rp}
}
Retrospective In-Context Learning for Temporal Credit Assignment with Large Language Models · NeurIPS 2025