← Search

Andrei Polubarov

4 accepted papers

2026

Decision Pre-Trained Transformer is a Scalable In-Context Reinforcement Learner

ICLR 2026poster

Recent progress in in-context reinforcement learning (ICRL) has demonstrated its potential for training generalist agents that can acquire new tasks directly at inference. Algorithm Distillation (AD) pioneered this paradigm and was subsequently scaled to multi-domain settings, although its ability t…

Cited by 0SourceScholar
2026

Object-Centric Latent Action Learning

AAAI 2026technical

Leveraging vast amounts of unlabeled internet video data for embodied AI is currently bottlenecked by the lack of action labels and the presence of action-correlated visual distractors. Although recent latent action policy optimization (LAPO) has shown promise in inferring proxy action labels from v

Cited by 0SourcePDFScholar
2025

Latent Action Learning Requires Supervision in the Presence of Distractors

ICML 2025poster

Recently, latent action learning, pioneered by Latent Action Policies (LAPO), have shown remarkable pre-training efficiency on observation-only data, offering potential for leveraging vast amounts of video available on the web for embodied AI. However, prior work has focused on distractor-free data,…

Cited by 35SourcePDFScholar
2025

Vintix: Action Model via In-Context Reinforcement Learning

ICML 2025poster

In-Context Reinforcement Learning (ICRL) represents a promising paradigm for developing generalist agents that learn at inference time through trial-and-error interactions, analogous to how large language models adapt contextually, but with a focus on reward maximization. However, the scalability of…