← Search

Damien Ernst

5 accepted papers

2026

Improving the Performance and Learning Stability of Parallelizable RNNs Designed for Ultra-Low Power Applications

ICML 2026spotlight

Sequence learning is dominated by Transformers and parallelizable recurrent neural networks such as state-space models, yet learning long-term dependencies remains challenging, and state-of-the-art designs trade power consumption for performance. The Bistable Memory Recurrent Unit (BMRU) was introdu…

Cited by 0SourceScholar
2026

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access

ICML 2026poster

Asymmetric actor-critic methods are widely used in partially observable reinforcement learning, but typically assume full state observability to condition the critic during training, which is often unrealistic in practice. We introduce the informed asymmetric actor-critic framework, allowing the cri…

Cited by 0SourceScholar
2025

A Theoretical Justification for Asymmetric Actor-Critic Algorithms

ICML 2025poster

In reinforcement learning for partially observable environments, many successful algorithms have been developed within the asymmetric learning paradigm. This paradigm leverages additional state information available at training time for faster learning. Although the proposed learning objectives are…

Cited by 1SourcePDFScholar
2023

IMP-MARL: a Suite of Environments for Large-scale Infrastructure Management Planning via MARL

NeurIPS 2023poster

We introduce IMP-MARL, an open-source suite of multi-agent reinforcement learning (MARL) environments for large-scale Infrastructure Management Planning (IMP), offering a platform for benchmarking the scalability of cooperative MARL methods in real-world engineering applications. In IMP, a multi-com…

2020

On Overfitting and Asymptotic Bias in Batch Reinforcement Learning with Partial Observability (Extended Abstract)

IJCAI 2020poster

When an agent has limited information on its environment, the suboptimality of an RL algorithm can be decomposed into the sum of two terms: a term related to an asymptotic bias (suboptimality with unlimited data) and a term due to overfitting (additional suboptimality due to limited data). In the co…

Cited by 0SourcePDFScholar