← Search

Todor Markov

3 accepted papers

2023

A Holistic Approach to Undesired Content Detection in the Real World

AAAI 2023technical

We present a holistic approach to building a robust and useful natural language classification system for real-world content moderation. The success of such a system relies on a chain of carefully designed and executed steps, including the design of content taxonomies and labeling instructions, data…

2020

Emergent Tool Use From Multi-Agent Autocurricula

ICLR 2020spotlight

Through multi-agent competition, the simple objective of hide-and-seek, and standard reinforcement learning algorithms at scale, we find that agents create a self-supervised autocurriculum inducing multiple distinct rounds of emergent strategy, many of which require sophisticated tool use and coordi…

Cited by 961SourcecodeScholar
2018

Best arm identification in multi-armed bandits with delayed feedback

AISTATS 2018poster

In this paper, we propose a generalization of the best arm identification problem in stochastic multi-armed bandits (MAB) to the setting where every pull of an arm is associated with delayed feedbacks. The delay in feedbacks increases the effective sample complexity of the algorithm, but can be offs…

Cited by 0SourcePDFScholar