2026
Reward Redistribution via Gaussian Process Likelihood Estimation
AAAI 2026technical
In many practical reinforcement learning tasks, feedback is only provided at the end of a long horizon, leading to sparse and delayed rewards. Existing reward redistribution methods typically assume that per-step rewards are independent, thus overlooking interdependencies among state–action pairs. I