← Search

David Meger

47 accepted papers

2026

Contractive Diffusion Policies: Robust Action Diffusion via Contractive Score-Based Sampling with Differential Equations

ICLR 2026poster

Diffusion policies have emerged as powerful generative models for offline policy learning, whose sampling process can be rigorously characterized by a score function guiding a Stochastic Differential Equation (SDE). However, the same score-based SDE modeling that grants diffusion policies the flexib…

Cited by 0SourceScholar
2026

VOCALoco: Viability-Optimized Cost-Aware Adaptive Locomotion

RA-L 2026

Recent advancements in legged robot locomotion have facilitated traversal over increasingly complex terrains. Despite this progress, many existing approaches rely on end-to-end deep reinforcement learning (DRL), which poses limitations in terms of safety and interpretability, especially when general

Cited by 1SourceScholar
2025

Convergence Theorems for Entropy-Regularized and Distributional Reinforcement Learning

NeurIPS 2025poster

In the pursuit of finding an optimal policy, reinforcement learning (RL) methods generally ignore the properties of learned policies apart from their expected return. Thus, even when successful, it is difficult to characterize which policies will be learned and what they will do. In this work, we pr…

Cited by 0SourceScholar
2025

Epistemic Uncertainty Estimation in Regression Ensemble Models with Pairwise Epistemic Estimators

NeurIPS 2025poster

This work introduces a novel approach, Pairwise Epistemic Estimators (PairEpEsts), for epistemic uncertainty estimation in ensemble models for regression tasks using pairwise-distance estimators (PaiDEs). By utilizing the pairwise distances between model components, PaiDEs establish bounds on entrop…

Cited by 0SourceScholar
2025

Generalizable Imitation Learning Through Pre-Trained Representations

ICRA 2025

In this paper, we leverage self-supervised vision transformer models and their emergent semantic abilities to improve the generalization abilities of imitation learning policies. We introduce DVK, an imitation learning algorithm that leverages rich pre-trained Visual Transformer patch-level embeddin

Cited by 5SourceScholar
2025

Learning Active Tactile Perception Through Belief-Space Control

ICRA 2025

Robots operating in an open world will encounter novel objects with unknown physical properties, such as mass, friction, or size. These robots will need to sense these properties through interaction prior to performing downstream tasks with the objects. We propose a method that autonomously learns t

Cited by 2SourceScholar
2025

Topological Mapping for Traversability-Aware Long-Range Navigation in Off-Road Terrain

ICRA 2025

Autonomous robots navigating in off-road terrain like forests open new opportunities for automation. While off-road navigation has been studied, existing work often relies on clearly delineated pathways. We present a method allowing for long-range planning, exploration and low-level control in unkno

Cited by 1SourceScholar
2024

Action Gaps and Advantages in Continuous-Time Distributional Reinforcement Learning

NeurIPS 2024poster

When decisions are made at high frequency, traditional reinforcement learning (RL) methods struggle to accurately estimate action values. In turn, their performance is inconsistent and often poor. Whether the performance of distributional RL (DRL) agents suffers similarly, however, is unknown. In th…

Cited by 1SourcePDFScholar
2024

Parseval Regularization for Continual Reinforcement Learning

NeurIPS 2024poster

Plasticity loss, trainability loss, and primacy bias have been identified as issues arising when training deep neural networks on sequences of tasks---referring to the increased difficulty in training on new tasks. We propose to use Parseval regularization, which maintains orthogonality of weight ma…

Cited by 0SourcePDFScholar
2024

Shedding Light on Large Generative Networks: Estimating Epistemic Uncertainty in Diffusion Models

UAI 2024poster

Generative diffusion models, notable for their large parameter count (exceeding 100 million) and operation within high-dimensional image spaces, pose significant challenges for traditional uncertainty estimation methods due to computational demands. In this work, we introduce an innovative framework…

2024

Uncertainty-aware hybrid paradigm of nonlinear MPC and model-based RL for offroad navigation: Exploration of transformers in the predictive model

ICRA 2024poster

In this paper, we investigate a hybrid scheme that combines nonlinear model predictive control (MPC) and model-based reinforcement learning (RL) for navigation planning of an autonomous model car across offroad, unstructured terrains without relying on predefined maps. Our innovative approach takes…

Cited by 4SourcecodeScholar
2023

ANSEL Photobot: A Robot Event Photographer with Semantic Intelligence

ICRA 2023poster

Our work examines the way in which large language models can be used for robotic planning and sampling in the context of automated photographic documentation. Specifically, we illustrate how to produce a photo-taking robot with an exceptional level of semantic awareness by leveraging recent advances…

Cited by 9SourceScholar
2023

For SALE: State-Action Representation Learning for Deep Reinforcement Learning

NeurIPS 2023poster

In reinforcement learning (RL), representation learning is a proven tool for complex image-based tasks, but is often overlooked for environments with low-level states, such as physical control problems. This paper introduces SALE, a novel approach for learning embeddings that model the nuanced inte…

2023

Hypernetworks for Zero-Shot Transfer in Reinforcement Learning

AAAI 2023technical

In this paper, hypernetworks are trained to generate behaviors across a range of unseen task conditions, via a novel TD-based training objective and data from a set of near-optimal RL solutions for training tasks. This work relates to meta RL, contextual RL, and transfer learning, with a particular…

Cited by 20SourcePDFScholar
2023

Normalizing Flow Ensembles for Rich Aleatoric and Epistemic Uncertainty Modeling

AAAI 2023technical

In this work, we demonstrate how to reliably estimate epistemic uncertainty while maintaining the flexibility needed to capture complicated aleatoric distributions. To this end, we propose an ensemble of Normalizing Flows (NF), which are state-of-the-art in modeling aleatoric uncertainty. The ensemb…

2022

Continuous MDP Homomorphisms and Homomorphic Policy Gradient

NeurIPS 2022accept

Abstraction has been widely studied as a way to improve the efficiency and generalization of reinforcement learning algorithms. In this paper, we study abstraction in the continuous-control setting. We extend the definition of MDP homomorphisms to encompass continuous actions in continuous state spa…

2022

Distributional Hamilton-Jacobi-Bellman Equations for Continuous-Time Reinforcement Learning

ICML 2022spotlight

Continuous-time reinforcement learning offers an appealing formalism for describing control problems in which the passage of time is not naturally divided into discrete increments. Here we consider the problem of predicting the distribution of returns obtained by an agent interacting in a continuous…

Cited by 14SourcePDFScholar
2022

Visuotactile-RL: Learning Multimodal Manipulation Policies with Deep Reinforcement Learning

ICRA 2022poster

Manipulating objects with dexterity requires timely feedback that simultaneously leverages the senses of vision and touch. In this paper, we focus on the problem setting where both visual and tactile sensors provide pixel-level feedback for Visuotactile reinforcement learning agents. We investigate…

Cited by 37SourceScholar
2022

Why Should I Trust You, Bellman? The Bellman Error is a Poor Replacement for Value Error

ICML 2022spotlight

In this work, we study the use of the Bellman equation as a surrogate objective for value prediction accuracy. While the Bellman equation is uniquely solved by the true value function over all state-action pairs, we find that the Bellman error (the difference between both sides of the equation) is a…

Cited by 41SourcePDFScholar
2021

A Deep Reinforcement Learning Approach to Marginalized Importance Sampling with the Successor Representation

ICML 2021spotlight

Marginalized importance sampling (MIS), which measures the density ratio between the state-action occupancy of a target policy and that of a sampling distribution, is a promising approach for off-policy evaluation. However, current state-of-the-art MIS methods rely on complex optimization tricks and…

2021

Active 3D Shape Reconstruction from Vision and Touch

NeurIPS 2021poster

Humans build 3D understandings of the world through active object exploration, using jointly their senses of vision and touch. However, in 3D shape reconstruction, most recent progress has relied on static datasets of limited sensory data such as RGB images, depth maps or haptic readings, leaving th…

2021

Latent Attention Augmentation for Robust Autonomous Driving Policies

IROS 2021poster

Model-free reinforcement learning has become a viable approach for vision-based robot control. However, sample complexity and adaptability to domain shifts remain persistent challenges when operating in high-dimensional observation spaces (images, LiDAR), such as those that are involved in autonomou…

Cited by 4SourceScholar
2021

Learning Intuitive Physics with Multimodal Generative Models

AAAI 2021technical

Predicting the future interaction of objects when they come into contact with their environment is key for autonomous agents to take intelligent and anticipatory actions. This paper presents a perception framework that fuses visual and tactile feedback to make predictions about the expected motion…

2021

Multimodal dynamics modeling for off-road autonomous vehicles

ICRA 2021poster

Dynamics modeling in outdoor and unstructured environments is difficult because different elements in the environment interact with the robot in ways that can be hard to predict. Leveraging multiple sensors to perceive maximal information about the robot’s environment is thus crucial when building a…

Cited by 17SourceScholar
2021

Trajectory-Constrained Deep Latent Visual Attention for Improved Local Planning in Presence of Heterogeneous Terrain

IROS 2021poster

We present a reward-predictive, model-based learning method featuring trajectory-constrained visual attention for use in mapless, local visual navigation tasks. Our method learns to place visual attention at locations in latent image space which follow trajectories caused by vehicle control actions…

Cited by 6SourceScholar
2020

3D Shape Reconstruction from Vision and Touch

NeurIPS 2020poster

When a toddler is presented a new toy, their instinctual behaviour is to pick it up and inspect it with their hand and eyes in tandem, clearly searching over its surface to properly understand what they are playing with. At any instance here, touch provides high fidelity localized information while…

2020

An Equivalence between Loss Functions and Non-Uniform Sampling in Experience Replay

NeurIPS 2020poster

Prioritized Experience Replay (PER) is a deep reinforcement learning technique in which agents learn from transitions sampled with non-uniform probability proportionate to their temporal-difference error. We show that any loss function evaluated with non-uniformly sampled data can be transformed int…

2020

Learning Domain Randomization Distributions for Training Robust Locomotion Policies

IROS 2020poster

This paper considers the problem of learning behaviors in simulation without knowledge of the precise dynamical properties of the target robot platform(s). In this context, our learning goal is to mutually maximize task efficacy on each environment considered and generalization across the widest pos…

Cited by 27SourceScholar
2020

Learning the Latent Space of Robot Dynamics for Cutting Interaction Inference

IROS 2020poster

Utilization of latent space to capture a lower-dimensional representation of a complex dynamics model is explored in this work. The targeted application is of a robotic manipulator executing a complex environment interaction task, in particular, cutting a wooden object. We train two flavours of Vari…

Cited by 6SourcecodeScholar
2020

Learning to Drive Off Road on Smooth Terrain in Unstructured Environments Using an On-Board Camera and Sparse Aerial Images

ICRA 2020poster

We present a method for learning to drive on smooth terrain while simultaneously avoiding collisions in challenging off-road and unstructured outdoor environments using only visual inputs. Our approach applies a hybrid model-based and model-free reinforcement learning method that is entirely self-su…

Cited by 45SourceScholar
2020

PresSense: Passive Respiration Sensing via Ambient WiFi Signals in Noisy Environments

IROS 2020poster

Passive sensing with ambient WiFi signals is a promising technique that will enable new types of human-robot interactions while preserving users' privacy. Here, we present PresSense, a system for human respiration sensing in noisy environments. Unlike existing WiFi-based respiration sensors, we empl…

Cited by 10SourceScholar
2020

Vision-Based Goal-Conditioned Policies for Underwater Navigation in the Presence of Obstacles

RSS 2020poster

We present Nav2Goal, a data-efficient and end-to-end learning method for goal-conditioned visual navigation. Our technique is used to train a navigation policy that enables a robot to navigate close to sparse geographic waypoints provided by a user without any prior map, all while avoiding obstacles…

Cited by 63SourcePDFScholar
2019

GEOMetrics: Exploiting Geometric Structure for Graph-Encoded Objects

ICML 2019oral

Mesh models are a promising approach for encoding the structure of 3D objects. Current mesh reconstruction systems predict uniformly distributed vertex locations of a predetermined graph through a series of graph convolutions, leading to compromises with respect to performance or resolution. In this…

2019

Uncertainty Aware Learning from Demonstrations in Multiple Contexts using Bayesian Neural Networks

ICRA 2019poster

Diversity of environments is a key challenge that causes learned robotic controllers to fail due to the discrepancies between the training and evaluation conditions. Training from demonstrations in various conditions can mitigate - but not completely prevent - such failures. Learned controllers such…

Cited by 24SourceScholar
2018

Cost Adaptation for Robust Decentralized Swarm Behaviour

IROS 2018poster

Decentralized receding horizon control (D-RHC) provides a mechanism for coordination in multiagent settings without a centralized command center. However, combining a set of different goals, costs, and constraints to form an efficient optimization objective for D-RHC can be difficult. To allay this…

Cited by 3SourcecodeScholar
2018

Multi-View Silhouette and Depth Decomposition for High Resolution 3D Object Representation

NeurIPS 2018poster

We consider the problem of scaling deep generative shape models to high-resolution. Drawing motivation from the canonical view representation of objects, we introduce a novel method for the fast up-sampling of 3D objects in voxel space through networks that perform super-resolution on the six orthog…

2018

Synthesizing Neural Network Controllers with Probabilistic Model-Based Reinforcement Learning

IROS 2018poster

We present an algorithm for rapidly learning neural network policies for robotics systems. The algorithm follows the model-based reinforcement learning paradigm and improves upon existing algorithms: PILeO and a sample-based version of PILeo with neural network dynamics (Deep-PILeO). To improve conv…

Cited by 52SourceScholar
2015

Learning legged swimming gaits from experience

ICRA 2015poster

We present an end-to-end framework for realizing fully automated gait learning for a complex underwater legged robot. Using this framework, we demonstrate that a hexapod flipper-propelled robot can learn task-specific control policies purely from experience data. Our method couples a state-of-the-ar…

Cited by 49SourceScholar