← Search

Michael L. Littman

7 accepted papers

2026

Different Usage of Shared Components Explains Behavioral Variance in LLMs

ICML 2026poster

One of the most common complaints about large language models (LLMs) is their prompt sensitivity---i.e., the fact that their ability to perform a task or provide a correct answer to a question can depend unpredictably on the way the question is posed. We investigate this variation by comparing two v…

Cited by 0SourceScholar
2022

On the (In)Tractability of Reinforcement Learning for LTL Objectives

IJCAI 2022poster

In recent years, researchers have made significant progress in devising reinforcement-learning algorithms for optimizing linear temporal logic (LTL) objectives and LTL-like objectives. Despite these advancements, there are fundamental limitations to how well this problem can be solved. Previous stu…

Cited by 22SourcePDFScholar
2022

On the Expressivity of Markov Reward (Extended Abstract)

IJCAI 2022poster

Reward is the driving force for reinforcement-learning agents. We here set out to understand the expressivity of Markov reward as a way to capture tasks that we would want an agent to perform. We frame this study around three new abstract notions of "task": (1) a set of acceptable behaviors…

Cited by 0SourcePDFScholar
2021

Deep Radial-Basis Value Functions for Continuous Control

AAAI 2021technical

A core operation in reinforcement learning (RL) is finding an action that is optimal with respect to a learned value function. This operation is often challenging when the learned value function takes continuous actions as input. We introduce deep radial-basis value functions (RBVFs): value function…

2021

Lipschitz Lifelong Reinforcement Learning

AAAI 2021technical

We consider the problem of knowledge transfer when an agent is facing a series of Reinforcement Learning (RL) tasks. We introduce a novel metric between Markov Decision Processes and establish that close MDPs have close optimal value functions. Formally, the optimal value functions are Lipschitz con…

2017

Interactive Learning from Policy-Dependent Human Feedback

ICML 2017poster

This paper investigates the problem of interactively learning behaviors communicated by a human teacher using positive and negative feedback. Much previous work on this problem has made the assumption that people provide feedback for decisions that is dependent on the behavior they are teaching and…

Cited by 387SourcePDFScholar