← Search

Yuri Burda

6 accepted papers

2024

Let's Verify Step by Step

ICLR 2024poster

In recent years, large language models have greatly improved in their ability to perform complex multi-step reasoning. However, even state-of-the-art models still regularly produce logical mistakes. To train more reliable models, we can turn either to outcome supervision, which provides feedback for…

2019

Large-Scale Study of Curiosity-Driven Learning

ICLR 2019poster

Reinforcement learning algorithms rely on carefully engineered rewards from the environment that are extrinsic to the agent. However, annotating each environment with hand-designed, dense rewards is difficult and not scalable, motivating the need for developing reward functions that are intrinsic to…

2018

Learning Policy Representations in Multiagent Systems

ICML 2018oral

Modeling agent behavior is central to understanding the emergence of complex phenomena in multiagent systems. Prior work in agent modeling has largely been task-specific and driven by hand-engineering domain-specific prior knowledge. We propose a general learning framework for modeling agent behavio…

Cited by 153SourcePDFScholar
2017

On the Quantitative Analysis of Decoder-Based Generative Models

ICLR 2017poster

The past several years have seen remarkable progress in generative models which produce convincing samples of images and other modalities. A shared component of some popular models such as generative adversarial networks and generative moment matching networks, is a decoder network, a parametric dee…

Cited by 285SourcecodeScholar
2015

Accurate and conservative estimates of MRF log-likelihood using reverse annealing

AISTATS 2015poster

Markov random fields (MRFs) are difficult to evaluate as generative models because computing the test log-probabilities requires the intractable partition function. Annealed importance sampling (AIS) is widely used to estimate MRF partition functions, and often yields quite accurate results. However…

Cited by 79SourcePDFScholar