← Search

Yichi Zhou

9 accepted papers

2024

Stabilizing Policy Gradients for Stochastic Differential Equations via Consistency with Perturbation Process

ICML 2024poster

Considering generating samples with high rewards, we focus on optimizing deep neural networks parameterized stochastic differential equations (SDEs), the advanced generative models with high expressiveness, with policy gradient, the leading algorithm in reinforcement learning. Nevertheless, when app…

Cited by 1SourcePDFScholar
2023

Efficiently incorporating quintuple interactions into geometric deep learning force fields

NeurIPS 2023poster

Machine learning force fields (MLFFs) have instigated a groundbreaking shift in molecular dynamics (MD) simulations across a wide range of fields, such as physics, chemistry, biology, and materials science. Incorporating higher order many-body interactions can enhance the expressiveness and accuracy…

2022

Simultaneously Learning Stochastic and Adversarial Bandits with General Graph Feedback

ICML 2022spotlight

The problem of online learning with graph feedback has been extensively studied in the literature due to its generality and potential to model various learning tasks. Existing works mainly study the adversarial and stochastic feedback separately. If the prior knowledge of the feedback mechanism is u…

Cited by 13SourcePDFScholar
2020

Lazy-CFR: fast and near-optimal regret minimization for extensive games with imperfect information

ICLR 2020poster

Counterfactual regret minimization (CFR) methods are effective for solving two-player zero-sum extensive games with imperfect information with state-of-the-art results. However, the vanilla CFR has to traverse the whole game tree in each round, which is time-consuming in large-scale games. In thi…

Cited by 15SourceScholar
2020

Posterior sampling for multi-agent reinforcement learning: solving extensive games with imperfect information

ICLR 2020talk

Posterior sampling for reinforcement learning (PSRL) is a useful framework for making decisions in an unknown environment. PSRL maintains a posterior distribution of the environment and then makes planning on the environment sampled from the posterior distribution. Though PSRL works well on single-…

Cited by 24SourceScholar
2018

Racing Thompson: an Efficient Algorithm for Thompson Sampling with Non-conjugate Priors

ICML 2018oral

Thompson sampling has impressive empirical performance for many multi-armed bandit problems. But current algorithms for Thompson sampling only work for the case of conjugate priors since they require to perform online Bayesian posterior inference, which is a difficult task when the prior is not conj…

Cited by 6SourcePDFScholar