← Search

Claudia Linnhoff-Popien

5 accepted papers

2023

Attention-Based Recurrence for Multi-Agent Reinforcement Learning under Stochastic Partial Observability

ICML 2023poster

Stochastic partial observability poses a major challenge for decentralized coordination in multi-agent reinforcement learning but is largely neglected in state-of-the-art research due to a strong focus on state-based centralized training for decentralized execution (CTDE) and benchmarks that lack su…

2023

CROP: Towards Distributional-Shift Robust Reinforcement Learning Using Compact Reshaped Observation Processing

IJCAI 2023poster

The safe application of reinforcement learning (RL) requires generalization from limited training data to unseen scenarios. Yet, fulfilling tasks under changing circumstances is a key challenge in RL. Current state-of-the-art approaches for generalization apply data augmentation techniques to increa…

2021

Resilient Multi-Agent Reinforcement Learning with Adversarial Value Decomposition

AAAI 2021technical

We focus on resilience in cooperative multi-agent systems, where agents can change their behavior due to udpates or failures of hardware and software components. Current state-of-the-art approaches to cooperative multi-agent reinforcement learning (MARL) have either focused on idealized settings wit…

2021

Stochastic Market Games

IJCAI 2021poster

Some of the most relevant future applications of multi-agent systems like autonomous driving or factories as a service display mixed-motive scenarios, where agents might have conflicting goals. In these settings agents are likely to learn undesirable outcomes in terms of cooperation under independen…

Cited by 6SourcePDFScholar
2021

VAST: Value Function Factorization with Variable Agent Sub-Teams

NeurIPS 2021poster

Value function factorization (VFF) is a popular approach to cooperative multi-agent reinforcement learning in order to learn local value functions from global rewards. However, state-of-the-art VFF is limited to a handful of agents in most domains. We hypothesize that this is due to the flat factori…