← Search

Benjamin Rosman

20 accepted papers

2026

Position: the Stochastic Parrot in the Coal Mine. Model Collapse is a Threat to Low-Resource Communities

ICML 2026spotlight

Model collapse, the degradation in performance that arises when generative models are trained on the outputs of prior models, is an increasing concern as artificially generated content proliferates. Related critiques of large language models have highlighted their tendency to reproduce frequent patt…

Cited by 0SourceScholar
2026

Unsupervised Hierarchical Skill Discovery

ICML 2026poster

We consider the problem of unsupervised skill segmentation and hierarchical structure discovery in reinforcement learning. While recent approaches have sought to segment trajectories into reusable skills or options, most rely on action labels, rewards, or handcrafted annotations, limiting their appl…

Cited by 0SourceScholar
2025

Make Haste Slowly: A Theory of Emergent Structured Mixed Selectivity in Feature Learning ReLU Networks

ICLR 2025poster

In spite of finite dimension ReLU neural networks being a consistent factor behind recent deep learning successes, a theory of feature learning in these models remains elusive. Currently, insightful theories still rely on assumptions including the linearity of the network computations, unstructured…

Cited by 1SourcePDFScholar
2025

The Esethu Framework: Reimagining Sustainable Dataset Governance and Curation for Low-Resource Languages

ACL 2025long

This paper presents the Esethu Framework, a sustainable data curation framework specifically designed to empower local communities and ensure equitable benefit-sharing from their linguistic resource. This framework is supported by the Esethu license, a novel community-centric data license. As a proo…

Cited by 0SourcePDFScholar
2024

Skill Machines: Temporal Logic Skill Composition in Reinforcement Learning

ICLR 2024poster

It is desirable for an agent to be able to solve a rich variety of problems that can be specified through language in the same environment. A popular approach towards obtaining such agents is to reuse skills learned in prior tasks to generalise compositionally to new ones. However, this is a challen…

2024

The Zeno’s Paradox of ‘Low-Resource’ Languages

EMNLP 2024main

The disparity in the languages commonly studied in Natural Language Processing (NLP) is typically reflected by referring to languages as low vs high-resourced. However, there is limited consensus on what exactly qualifies as a ‘low-resource language.’ To understand how NLP papers define and study ‘l…

2023

Dynamics Generalisation in Reinforcement Learning via Adaptive Context-Aware Policies

NeurIPS 2023poster

While reinforcement learning has achieved remarkable successes in several domains, its real-world application is limited due to many methods failing to generalise to unfamiliar conditions. In this work, we consider the problem of generalising to new transition dynamics, corresponding to cases in whi…

2022

Autonomous Learning of Object-Centric Abstractions for High-Level Planning

ICLR 2022poster

We propose a method for autonomously learning an object-centric representation of a continuous and high-dimensional environment that is suitable for planning. Such representations can immediately be transferred between tasks that share the same types of objects, resulting in agents that require fewe…

Cited by 28SourcePDFScholar
2022

Generalisation in Lifelong Reinforcement Learning through Logical Composition

ICLR 2022poster

We leverage logical composition in reinforcement learning to create a framework that enables an agent to autonomously determine whether a new task can be immediately solved using its existing abilities, or whether a task-specific skill should be learned. In the latter case, the proposed algorithm al…

Cited by 26SourcePDFScholar
2019

Composing Value Functions in Reinforcement Learning

ICML 2019oral

An important property for lifelong-learning agents is the ability to combine existing skills to solve new unseen tasks. In general, however, it is unclear how to compose existing skills in a principled manner. Under the assumption of deterministic dynamics, we prove that optimal value function compo…

Cited by 63SourcePDFScholar
2018

Hierarchical Subtask Discovery with Non-Negative Matrix Factorization

ICLR 2018poster

Hierarchical reinforcement learning methods offer a powerful means of planning flexible behavior in complicated domains. However, learning an appropriate hierarchical decomposition of a domain into subtasks remains a substantial challenge. We present a novel algorithm for subtask discovery, based on…

Cited by 12SourcePDFScholar
2018

Real-Time Motion Planning in Changing Environments Using Topology-Based Encoding of Past Knowledge

IROS 2018poster

Trajectory planning and replanning in complex environments often reuses very little information from the previous solutions. This is particularly evident when the motion is repeated multiple times with only a limited amount of variation between each run. To address this issue, we propose the DRM-con…

Cited by 5SourceScholar
2018

Zero-Shot Transfer with Deictic Object-Oriented Representation in Reinforcement Learning

NeurIPS 2018poster

Object-oriented representations in reinforcement learning have shown promise in transfer learning, with previous research introducing a propositional object-oriented framework that has provably efficient learning bounds with respect to sample complexity. However, this framework has limitations in te…

Cited by 17SourcePDFScholar
2016

Trajectory learning from human demonstrations via manifold mapping

IROS 2016poster

This work proposes a framework that enables arbitrary robots with unknown kinematics models to imitate human demonstrations to acquire a skill, and reproduce it in real-time. The diversity of robots active in non-laboratory environments is growing constantly, and to this end we present an approach f…

Cited by 11SourceScholar
2015

Nonparametric Bayesian reward segmentation for skill discovery using inverse reinforcement learning

IROS 2015poster

We present a method for segmenting a set of unstructured demonstration trajectories to discover reusable skills using inverse reinforcement learning (IRL). Each skill is characterised by a latent reward function which the demonstrator is assumed to be optimizing. The skill boundaries and the number…

Cited by 88SourceScholar