← Search

Martin Mladenov

6 accepted papers

2024

Demystifying Embedding Spaces using Large Language Models

ICLR 2024poster

Embeddings have become a pivotal means to represent complex, multi-faceted information about entities, concepts, and relationships in a condensed and useful format. Nevertheless, they often preclude direct interpretation. While downstream tasks make use of these compressed representations, meaningfu…

Cited by 18SourcePDFScholar
2024

Recommender Ecosystems: A Mechanism Design Perspective on Holistic Modeling and Optimization

AAAI 2024technical

Modern recommender systems lie at the heart of complex recommender ecosystems that couple the behavior of users, content providers, vendors, advertisers, and other actors. Despite this, the focus of much recommender systems research and deployment is on the local, myopic optimization of the recommen…

Cited by 2SourcePDFScholar
2023

Reinforcement Learning with History Dependent Dynamic Contexts

ICML 2023poster

We introduce *Dynamic Contextual Markov Decision Processes (DCMDPs)*, a novel reinforcement learning framework for history-dependent environments that generalizes the contextual MDP framework to handle non-Markov environments, where contexts change over time. We consider special cases of the model,…

Cited by 10SourcePDFScholar
2021

Meta-Thompson Sampling

ICML 2021spotlight

Efficient exploration in bandits is a fundamental online learning problem. We propose a variant of Thompson sampling that learns to explore better as it interacts with bandit instances drawn from an unknown prior. The algorithm meta-learns the prior and thus we call it MetaTS. We propose several eff…

Cited by 84SourcePDFScholar
2020

Differentiable Meta-Learning of Bandit Policies

NeurIPS 2020poster

Exploration policies in Bayesian bandits maximize the average reward over problem instances drawn from some distribution P. In this work, we learn such policies for an unknown distribution P using samples from P. Our approach is a form of meta-learning and exploits properties of P without making str…

2020

Optimizing Long-term Social Welfare in Recommender Systems: A Constrained Matching Approach

ICML 2020poster

Most recommender systems (RS) research assumes that a user’s utility can be maximized independently of the utility of the other agents (e.g., other users, content providers). In realistic settings, this is often not true – the dynamics of an RS ecosystem couple the long-term utility of all agents. I…

Cited by 74SourcePDFScholar