← Search

Renos Zabounidis

2 accepted papers

2026

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning

ICML 2026poster

We propose Re-FORC, an adaptive reward prediction method that, given a context, enables prediction of the expected future rewards as a function of the number of future thinking tokens. Re-FORC trains a lightweight adapter on reasoning models, demonstrating improved prediction with longer reasoning a…

Cited by 0SourceScholar
2022

Concept Learning for Interpretable Multi-Agent Reinforcement Learning

CoRL 2022poster

Multi-agent robotic systems are increasingly operating in real-world environments in close proximity to humans, yet are largely controlled by policy models with inscrutable deep neural network representations. We introduce a method for incorporating interpretable concepts from a domain expert into m…

Cited by 23SourceScholar