← Search

Sanjit Seshia

6 accepted papers

2026

Automata-Conditioned Cooperative Multi-Agent Reinforcement Learning

ICML 2026poster

We study learning multi-task, multi-agent policies for cooperative, temporal objectives, under centralized training, decentralized execution. In this setting, using automata to represent tasks assigned to agents enables breaking down a team-level objective into simpler, smaller sub-tasks. However, e…

Cited by 0SourceScholar
2021

Entropy-Guided Control Improvisation

RSS 2021poster

High level declarative constraints provide a powerful (and popular) way to define and construct control policies; however; most synthesis algorithms do not support specifying the degree of randomness (unpredictability) of the resulting controller. In many contexts; e.g.; patrolling; testing; behavio…

Cited by 5SourcePDFScholar
2020

Learning Heuristics for Quantified Boolean Formulas through Reinforcement Learning

ICLR 2020poster

We demonstrate how to learn efficient heuristics for automated reasoning algorithms for quantified Boolean formulas through deep reinforcement learning. We focus on a backtracking search algorithm, which can already solve formulas of impressive size - up to hundreds of thousands of variables. The ma…

Cited by 41SourceScholar
2019

On the Utility of Learning about Humans for Human-AI Coordination

NeurIPS 2019poster

While we would like agents that can coordinate with humans, current algorithms such as self-play and population-based training create agents that can coordinate with themselves. Agents that assume their partner to be optimal or similar to them can converge to coordination protocols that fail to unde…

2018

Learning Task Specifications from Demonstrations

NeurIPS 2018poster

Real-world applications often naturally decompose into several sub-tasks. In many settings (e.g., robotics) demonstrations provide a natural way to specify the sub-tasks. However, most methods for learning from demonstrations either do not provide guarantees that the artifacts learned for th…

Cited by 96SourcePDFScholar