← Search

Rasul Tutunov

6 accepted papers

2026

Scalable Power Sampling: Unlocking Efficient, Training-Free Reasoning for LLMs via Distribution Sharpening

ICML 2026poster

Reinforcement learning (RL) post-training is a dominant approach for improving the reasoning performance of large language models (LLMs), yet growing evidence suggests that its gains arise primarily from distribution sharpening rather than the acquisition of new capabilities. Recent work has shown t…

Cited by 0SourceScholar
2023

Online PCA in Converging Self-consistent Field Equations

NeurIPS 2023poster

Self-consistent Field (SCF) equation is a type of nonlinear eigenvalue problem in which the matrix to be eigen-decomposed is a function of its own eigenvectors. It is of great significance in computational science for its connection to the Schrödinger equation. Traditional fixed-point iteration meth…

Cited by 0SourcePDFScholar
2022

Optimistic Tree Searches for Combinatorial Black-Box Optimization

NeurIPS 2022accept

The optimization of combinatorial black-box functions is pervasive in computer science and engineering. However, the combinatorial explosion of the search space and lack of natural ordering pose significant challenges for current techniques from a theoretical and practical perspective, and require n…

Cited by 3SourcePDFScholar
2018

Distributed Multitask Reinforcement Learning with Quadratic Convergence

NeurIPS 2018poster

Multitask reinforcement learning (MTRL) suffers from scalability issues when the number of tasks or trajectories grows large. The main reason behind this drawback is the reliance on centeralised solutions. Recent methods exploited the connection between MTRL and general consensus to propose scalable…

Cited by 13SourcePDFScholar
2015

Safe Policy Search for Lifelong Reinforcement Learning with Sublinear Regret

ICML 2015poster

Lifelong reinforcement learning provides a promising framework for developing versatile agents that can accumulate knowledge over a lifetime of experience and rapidly learn new tasks by building upon prior knowledge. However, current lifelong learning methods exhibit non-vanishing regret as the amou…

Cited by 88SourcePDFScholar