← Search

Osman Yagan

7 accepted papers

2026

Achieving Logarithmic Regret in KL-Regularized Zero-Sum Markov Games

ICML 2026poster

Reverse Kullback–Leibler (KL) divergence-based regularization with respect to a fixed reference policy is widely used in modern reinforcement learning to preserve the desired traits of the reference policy and sometimes to promote exploration (using uniform reference policy, known as entropy regular…

Cited by 0SourceScholar
2025

FedSPD: A Soft-clustering Approach for Personalized Decentralized Federated Learning

UAI 2025

Federated learning has recently gained popularity as a framework for distributed clients to collaboratively train a machine learning model using local data. While traditional federated learning relies on a central server for model aggregation, recent advancements adopt a decentralized framework, ena

Cited by 0SourcePDFScholar
2025

Pairwise Elimination with Instance-Dependent Guarantees for Bandits with Cost Subsidy

ICLR 2025poster

Multi-armed bandits (MAB) are commonly used in sequential online decision-making when the reward of each decision is an unknown random variable. In practice however, the typical goal of maximizing total reward may be less important than minimizing the total cost of the decisions taken, subject to a…

Cited by 0SourcePDFScholar
2021

A Unified Approach to Translate Classical Bandit Algorithms to Structured Bandits

ICASSP 2021accepted

We consider a finite-armed structured bandit problem in which mean rewards of different arms are known functions of a common hidden parameter θ <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">*</sup> . This problem setting subsumes several previously st…

Cited by 0SourceScholar
2021

Leveraging A Multiple-Strain Model with Mutations in Analyzing the Spread of Covid-19

ICASSP 2021accepted

The spread of COVID-19 has been among the most devastating events affecting the health and well-being of humans worldwide since World War II. A key scientific goal concerning COVID-19 is to develop mathematical models that help us to understand and predict its spreading behavior, as well as to provi…

Cited by 0SourceScholar