← Search

Kenshi Abe

19 accepted papers

2026

Asymmetric Perturbation in Solving Bilinear Saddle-Point Optimization

ICML 2026oral

This paper proposes an asymmetric perturbation technique for solving bilinear saddle-point optimization problems, commonly arising in minimax problems, game theory, and constrained optimization. Perturbing payoffs or values is known to be effective in stabilizing learning dynamics and equilibrium co…

Cited by 0SourceScholar
2026

Last-Iterate Convergence of Regularized Gradient Methods for Stochastic Monotone Variational Inequalities

ICML 2026poster

We study last-iterate convergence for stochastic smooth and monotone variational inequalities (VIs), a framework that captures convex-concave saddle points and Nash equilibrium computation in monotone games with noisy payoff feedback. In contrast to the well-understood average-iterate guarantees, an…

Cited by 0SourceScholar
2025

Boosting Perturbed Gradient Ascent for Last-Iterate Convergence in Games

ICLR 2025poster

This paper presents a payoff perturbation technique, introducing a strong convexity to players' payoff functions in games. This technique is specifically designed for first-order methods to achieve last-iterate convergence in games where the gradient of the payoff functions is monotone in the strate…

Cited by 0SourcePDFScholar
2025

Regularized Best-of-N Sampling with Minimum Bayes Risk Objective for Language Model Alignment

NAACL 2025long

Best-of-N (BoN) sampling with a reward model has been shown to be an effective strategy for aligning Large Language Models (LLMs) to human preferences at the time of decoding. BoN sampling is susceptible to a problem known as reward hacking when the accuracy of the reward model is not high enough. B…

2025

Synchronization in Learning in Periodic Zero-Sum Games Triggers Divergence from Nash Equilibrium

AAAI 2025technical

Learning in zero-sum games studies a situation where multiple agents competitively learn their strategy. In such multi-agent learning, we often see that the strategies cycle around their optimum, i.e., Nash equilibrium. When a game periodically varies (called a ``periodic'' game), however, the Nash…

Cited by 0SourcePDFScholar
2024

Adaptively Perturbed Mirror Descent for Learning in Games

ICML 2024poster

This paper proposes a payoff perturbation technique for the Mirror Descent (MD) algorithm in games where the gradient of the payoff functions is monotone in the strategy profile space, potentially containing additive noise. The optimistic family of learning algorithms, exemplified by optimistic MD,…

2024

Filtered Direct Preference Optimization

EMNLP 2024main

Reinforcement learning from human feedback (RLHF) plays a crucial role in aligning language models with human preferences. While the significance of dataset quality is generally recognized, explicit investigations into its impact within the RLHF framework, to our knowledge, have been limited. This p…

2024

Memory Asymmetry Creates Heteroclinic Orbits to Nash Equilibrium in Learning in Zero-Sum Games

AAAI 2024technical

Learning in games considers how multiple agents maximize their own rewards through repeated games. Memory, an ability that an agent changes his/her action depending on the history of actions in previous games, is often introduced into learning to explore more clever strategies and discuss the decisi…

2024

Model-Based Minimum Bayes Risk Decoding for Text Generation

ICML 2024poster

Minimum Bayes Risk (MBR) decoding has been shown to be a powerful alternative to beam search decoding in a variety of text generation tasks. MBR decoding selects a hypothesis from a pool of hypotheses that has the least expected risk under a probability model according to a given utility function. S…

2023

Last-Iterate Convergence with Full and Noisy Feedback in Two-Player Zero-Sum Games

AISTATS 2023poster

This paper proposes Mutation-Driven Multiplicative Weights Update (M2WU) for learning an equilibrium in two-player zero-sum normal-form games and proves that it exhibits the last-iterate convergence property in both full and noisy feedback settings. In the former, players observe their exact gradien…

2023

Learning in Multi-Memory Games Triggers Complex Dynamics Diverging from Nash Equilibrium

IJCAI 2023poster

Repeated games consider a situation where multiple agents are motivated by their independent rewards throughout learning. In general, the dynamics of their learning become complex. Especially when their rewards compete with each other like zero-sum games, the dynamics often do not converge to their…

2022

Anytime Capacity Expansion in Medical Residency Match by Monte Carlo Tree Search

IJCAI 2022poster

This paper considers the capacity expansion problem in two-sided matchings, where the policymaker is allowed to allocate some extra seats as well as the standard seats. In medical residency match, each hospital accepts a limited number of doctors. Such capacity constraints are typically given in adv…

2022

Mutation-driven follow the regularized leader for last-iterate convergence in zero-sum games

UAI 2022poster

In this study, we consider a variant of the Follow the Regularized Leader (FTRL) dynamics in two-player zero-sum games. FTRL is guaranteed to converge to a Nash equilibrium when time-averaging the strategies, while a lot of variants suffer from the issue of limit cycling behavior, i.e., lack the las…