← Search

Matthew Riemer

18 accepted papers

2026

Position: Collusion Risks Among AI Reasoning Agents Justify Certification Requirements for Making Market Decisions

ICML 2026poster

This position paper argues that AI agents with chain-of-thought reasoning capabilities are predisposed to exhibit collusive behavior and should be required to obtain behavioral certification before making decisions that affect economic markets. This is because integrating these agents into society c…

Cited by 0SourceScholar
2026

The Shepherd Test: How Will Super Intelligent Agents Balance Care and Control in Asymmetric Relationships?

AAAI 2026technical

This paper introduces the Shepherd Test, a new conceptual test for assessing the moral and relational dimensions of superintelligent artificial agents. The test is inspired by human interactions with animals, where ethical considerations about care, manipulation, and consumption arise in contexts of

Cited by 0SourcePDFScholar
2025

Combining Domain and Alignment Vectors Provides Better Knowledge-Safety Trade-offs in LLMs

ACL 2025short

There is a growing interest in training domain-expert LLMs that excel in specific technical fields compared to their general-purpose instruction-tuned counterparts. However, these expert models are not either explicitly trained to be safe, or experience a loss in their safety abilities in the proces…

2025

Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference

ICLR 2025poster

Realtime environments change even as agents perform action inference and learning, thus requiring high interaction frequencies to effectively minimize regret. However, recent advances in machine learning involve larger neural networks with longer inference times, raising questions about their applic…

2025

EpMAN: Episodic Memory AttentioN for Generalizing to Longer Contexts

ACL 2025long

Recent advances in Large Language Models (LLMs) have yielded impressive successes on many language tasks. However, efficient processing of long contexts using LLMs remains a significant challenge. We introduce **EpMAN** – a method for processing long contexts in an episodic memory module while holis…

2025

Handling Delay in Real-Time Reinforcement Learning

ICLR 2025poster

Real-time reinforcement learning (RL) introduces several challenges. First, policies are constrained to a fixed number of actions per second due to hardware limitations. Second, the environment may change while the network is still computing an action, leading to observational delay. The first issue…

2025

Position: Theory of Mind Benchmarks are Broken for Large Language Models

ICML 2025poster

Our paper argues that the majority of theory of mind benchmarks are broken because of their inability to directly test how large language models (LLMs) adapt to new partners. This problem stems from the fact that theory of mind benchmarks for LLMs are overwhelmingly inspired by the methods used to t…

Cited by 0SourcePDFScholar
2024

A Deep Dive into the Trade-Offs of Parameter-Efficient Preference Alignment Techniques

ACL 2024long

Large language models are first pre-trained on trillions of tokens and then instruction-tuned or aligned to specific preferences. While pre-training remains out of reach for most researchers due to the compute required, fine-tuning has become affordable thanks to parameter-efficient methods such as…

2024

Balancing Context Length and Mixing Times for Reinforcement Learning at Scale

NeurIPS 2024poster

Due to the recent remarkable advances in artificial intelligence, researchers have begun to consider challenging learning problems such as learning to generalize behavior from large offline datasets or learning online in non-Markovian environments. Meanwhile, recent advances in both of these areas h…

Cited by 3SourcePDFScholar
2024

ComVas: Contextual Moral Values Alignment System

IJCAI 2024poster

In contemporary society, the integration of artificial intelligence (AI) systems into various aspects of daily life raises significant ethical concerns. One critical aspect is to ensure that AI systems align with the moral values of the endusers. To that end, we introduce the Contextual Moral Value…

2022

Context-Specific Representation Abstraction for Deep Option Learning

AAAI 2022technical

Hierarchical reinforcement learning has focused on discovering temporally extended actions, such as options, that can provide benefits in problems requiring extensive exploration. One promising approach that learns these options end-to-end is the option-critic (OC) framework. We examine and show in…

2022

Continual Learning In Environments With Polynomial Mixing Times

NeurIPS 2022accept

The mixing time of the Markov chain induced by a policy limits performance in real-world continual learning scenarios. Yet, the effect of mixing times on learning in continual reinforcement learning (RL) remains underexplored. In this paper, we characterize problems that are of long-term interest to…

2022

Influencing Long-Term Behavior in Multiagent Reinforcement Learning

NeurIPS 2022accept

The main challenge of multiagent reinforcement learning is the difficulty of learning useful policies in the presence of other simultaneously learning agents whose changing behaviors jointly affect the environment's transition and reward dynamics. An effective approach that has recently emerged for…

2021

Efficient Black-Box Planning Using Macro-Actions with Focused Effects

IJCAI 2021poster

The difficulty of deterministic planning increases exponentially with search-tree depth. Black-box planning presents an even greater challenge, since planners must operate without an explicit model of the domain. Heuristics can make search more efficient, but goal-aware heuristics for black-box plan…

2019

Learning to Learn without Forgetting by Maximizing Transfer and Minimizing Interference

ICLR 2019poster

Lack of performance when it comes to continual learning over non-stationary distributions of data remains a major challenge in scaling neural network learning to more human realistic settings. In this work we propose a new conceptualization of the continual learning problem in terms of a temporally…

2018

Routing Networks: Adaptive Selection of Non-Linear Functions for Multi-Task Learning

ICLR 2018poster

Multi-task learning (MTL) with neural networks leverages commonalities in tasks to improve performance, but often suffers from task interference which reduces the benefits of transfer. To address this issue we introduce the routing network paradigm, a novel neural network and training algorithm. A r…

Cited by 302SourcePDFScholar
2016

Correcting Forecasts with Multifactor Neural Attention

ICML 2016poster

Automatic forecasting of time series data is a challenging problem in many industries. Current forecast models adopted by businesses do not provide adequate means for including data representing external factors that may have a significant impact on the time series, such as weather, national events,…

Cited by 43SourcePDFScholar