← Search

Georg Martius

62 accepted papers

2026

Differentiable Simulation of Hard Contacts with Soft Gradients for Learning and Control

ICLR 2026poster

Contact forces introduce discontinuities into robot dynamics that severely limit the use of simulators for gradient-based optimization. Penalty-based simulators such as MuJoCo, soften contact resolution to enable gradient computation. However, realistically simulating hard contacts requires stiff so…

Cited by 0SourceScholar
2026

SoftJAX & SoftTorch: Empowering Automatic Differentiation Libraries with Informative Gradients

ICML 2026oral

Automatic differentiation (AD) frameworks such as JAX and PyTorch have enabled gradient-based optimization for a wide range of scientific fields. Yet, many ''hard'' primitives in these libraries such as thresholding, Boolean logic, discrete indexing, and sorting operations yield zero or undefined gr…

Cited by 0SourceScholar
2026

Test-time Offline Reinforcement Learning on Goal-related Experience

ICML 2026poster

Foundation models compress a large amount of information in a single, large neural network, which can then be queried for individual tasks. There are strong parallels between this widespread framework and offline goal-conditioned reinforcement learning algorithms: a universal value function is train…

Cited by 0SourceScholar
2025

A Smooth Analytical Formulation of Collision Detection and Rigid Body Dynamics With Contact

IROS 2025

Generating intelligent robot behavior in contact-rich settings is a research problem where zeroth-order methods currently prevail. A major contributor to the success of such methods is their robustness in the face of non-smooth and discontinuous optimization landscapes that are characteristic of con

Cited by 5SourceScholar
2025

Advancing Out-of-Distribution Detection via Local Neuroplasticity

ICLR 2025poster

In the domain of machine learning, the assumption that training and test data share the same distribution is often violated in real-world scenarios, requiring effective out-of-distribution (OOD) detection. This paper presents a novel OOD detection method that leverages the unique local neuroplastic…

2025

Forecasting in Offline Reinforcement Learning for Non-stationary Environments

NeurIPS 2025spotlight

Offline Reinforcement Learning (RL) provides a promising avenue for training policies from pre-collected datasets when gathering additional interaction data is infeasible. However, existing offline RL methods often assume stationarity or only consider synthetic perturbations at test time—assumptions…

Cited by 0SourceScholar
2025

On the Transfer of Object-Centric Representation Learning

ICLR 2025poster

The goal of object-centric representation learning is to decompose visual scenes into a structured representation that isolates the entities into individual vectors. Recent successes have shown that object-centric representation learning can be scaled to real-world scenes by utilizing features from…

Cited by 1SourcePDFScholar
2025

SENSEI: Semantic Exploration Guided by Foundation Models to Learn Versatile World Models

ICML 2025poster

Exploration is a cornerstone of reinforcement learning (RL). Intrinsic motivation attempts to decouple exploration from external, task-based rewards. However, established approaches to intrinsic motivation that follow general principles such as information gain, often only uncover low-level interact…

Cited by 2SourcePDFScholar
2025

Temporally Consistent Object-Centric Learning by Contrasting Slots

CVPR 2025poster

Unsupervised object-centric learning from videos is a promising approach to extract structured representations from large, unlabeled collections of videos. To support downstream tasks like autonomous control, these representations must be both compositional and temporally consistent. Existing approa…

Cited by 1SourcePDFScholar
2025

Zero-Shot Offline Imitation Learning via Optimal Transport

ICML 2025poster

Zero-shot imitation learning algorithms hold the promise of reproducing unseen behavior from as little as a single demonstration at test time. Existing practical approaches view the expert demonstration as a sequence of goals, enabling imitation with a high-level goal selector, and a low-level goal-…

2024

Causal Action Influence Aware Counterfactual Data Augmentation

ICML 2024poster

Offline data are both valuable and practical resources for teaching robots complex behaviors. Ideally, learning agents should not be constrained by the scarcity of available demonstrations, but rather generalize beyond the training distribution. However, the complexity of real-world scenarios typica…

2024

Colored Noise in PPO: Improved Exploration and Performance through Correlated Action Sampling

AAAI 2024technical

Proximal Policy Optimization (PPO), a popular on-policy deep reinforcement learning method, employs a stochastic policy for exploration. In this paper, we propose a colored noise-based stochastic policy variant of PPO. Previous research highlighted the importance of temporal correlation in action no…

2024

Emergent mechanisms for long timescales depend on training curriculum and affect performance in memory tasks

ICLR 2024poster

Recurrent neural networks (RNNs) in the brain and \emph{in silico} excel at solving tasks with intricate temporal dependencies. Long timescales required for solving such tasks can arise from properties of individual neurons (single-neuron timescale, $\tau$, e.g., membrane time constant in biological…

2024

Identifying Terrain Physical Parameters From Vision - Towards Physical-Parameter-Aware Locomotion and Navigation

RA-L 2024

Identifying the physical properties of the surrounding environment is essential for robotic locomotion and navigation to deal with non-geometric hazards, such as slippery and deformable terrains. It would be of great benefit for robots to anticipate these extreme physical properties before contact;

Cited by 29SourceScholar
2024

LPGD: A General Framework for Backpropagation through Embedded Optimization Layers

ICML 2024poster

Embedding parameterized optimization problems as layers into machine learning architectures serves as a powerful inductive bias. Training such architectures with stochastic gradient descent requires care, as degenerate derivatives of the embedded optimization problem often render the gradients uninf…

2024

Learning Diverse Skills for Local Navigation under Multi-constraint Optimality

ICRA 2024poster

Despite many successful applications of data-driven control in robotics, extracting meaningful diverse behaviors remains a challenge. Typically, task performance needs to be compromised in order to achieve diversity. In many scenarios, task requirements are specified as a multitude of reward terms,…

Cited by 7SourceScholar
2024

Learning Hierarchical World Models with Adaptive Temporal Abstractions from Discrete Latent Dynamics

ICLR 2024spotlight

Hierarchical world models can significantly improve model-based reinforcement learning (MBRL) and planning by enabling reasoning across multiple time scales. Nonetheless, the majority of state-of-the-art MBRL methods employ flat, non-hierarchical models. We propose Temporal Hierarchies from Invarian…

Cited by 16SourcePDFScholar
2024

Learning with 3D rotations, a hitchhiker's guide to SO(3)

ICML 2024poster

Many settings in machine learning require the selection of a rotation representation. However, choosing a suitable representation from the many available options is challenging. This paper acts as a survey and guide through rotation representations. We walk through their properties that harm or bene…

2024

Modelling Microbial Communities with Graph Neural Networks

ICML 2024poster

Understanding the interactions and interplay of microorganisms is a great challenge with many applications in medical and environmental settings. In this work, we model bacterial communities directly from their genomes using graph neural networks (GNNs). GNNs leverage the inductive bias induced by t…

Cited by 2SourcePDFScholar
2024

Multi-View Causal Representation Learning with Partial Observability

ICLR 2024spotlight

We present a unified framework for studying the identifiability of representations learned from simultaneously observed views, such as different data modalities. We allow a partially observed setting in which each view constitutes a nonlinear mixture of a subset of underlying latent variables, which…

2024

The Expressive Leaky Memory Neuron: an Efficient and Expressive Phenomenological Neuron Model Can Solve Long-Horizon Tasks.

ICLR 2024poster

Biological cortical neurons are remarkably sophisticated computational devices, temporally integrating their vast synaptic input over an intricate dendritic tree, subject to complex, nonlinearly interacting internal biological processes. A recent study proposed to characterize this complexity by fi…

2023

Backpropagation through Combinatorial Algorithms: Identity with Projection Works

ICLR 2023poster

Embedding discrete solvers as differentiable layers has given modern deep learning architectures combinatorial expressivity and discrete reasoning capabilities. The derivative of these solvers is zero or undefined, therefore a meaningful replacement is crucial for effective gradient-based learning.…

2023

Benchmarking Offline Reinforcement Learning on Real-Robot Hardware

ICLR 2023top-25%

Learning policies from previously recorded data is a promising direction for real-world robotics tasks, as online learning is often infeasible. Dexterous manipulation in particular remains an open problem in its general form. The combination of offline reinforcement learning with large diverse datas…

Cited by 38SourcePDFScholar
2023

DEP-RL: Embodied Exploration for Reinforcement Learning in Overactuated and Musculoskeletal Systems

ICLR 2023top-25%

Muscle-actuated organisms are capable of learning an unparalleled diversity of dexterous movements despite their vast amount of muscles. Reinforcement learning (RL) on large musculoskeletal models, however, has not been able to show similar performance. We conjecture that ineffective exploration…

2023

Efficient Learning of High Level Plans from Play

ICRA 2023poster

Real-world robotic manipulation tasks remain an elusive challenge, since they involve both fine-grained environment interaction, as well as the ability to plan for long-horizon goals. Although deep reinforcement learning (RL) methods have shown encouraging results when planning end-to-end in high-di…

Cited by 5SourceScholar
2023

Object-Centric Learning for Real-World Videos by Predicting Temporal Feature Similarities

NeurIPS 2023poster

Unsupervised video-based object-centric learning is a promising avenue to learn structured representations from large, unlabeled video collections, but previous approaches have only managed to scale to real-world datasets in restricted domains. Recently, it was shown that the reconstruction of pre-t…

2023

Pink Noise Is All You Need: Colored Noise Exploration in Deep Reinforcement Learning

ICLR 2023top-25%

In off-policy deep reinforcement learning with continuous action spaces, exploration is often implemented by injecting action noise into the action selection process. Popular algorithms based on stochastic policies, such as SAC or MPO, inject white noise by sampling actions from uncorrelated Gaussia…

Cited by 51SourcePDFScholar
2023

Versatile Skill Control via Self-supervised Adversarial Imitation of Unlabeled Mixed Motions

ICRA 2023poster

Learning diverse skills is one of the main challenges in robotics. To this end, imitation learning approaches have achieved impressive results. These methods require explicitly labeled datasets or assume consistent skill execution to enable learning and active control of individual behaviors, which…

Cited by 34SourceScholar
2022

Curious Exploration via Structured World Models Yields Zero-Shot Object Manipulation

NeurIPS 2022accept

It has been a long-standing dream to design artificial agents that explore their environment efficiently via intrinsic motivation, similar to how children perform curious free play. Despite recent advances in intrinsically motivated reinforcement learning (RL), sample-efficient exploration in object…

Cited by 30SourcePDFScholar
2022

Embrace the Gap: VAEs Perform Independent Mechanism Analysis

NeurIPS 2022accept

Variational autoencoders (VAEs) are a popular framework for modeling complex data distributions; they can be efficiently trained via variational inference by maximizing the evidence lower bound (ELBO), at the expense of a gap to the exact (log-)marginal likelihood. While VAEs are commonly used for r…

2022

Learning Agile Skills via Adversarial Imitation of Rough Partial Demonstrations

CoRL 2022oral

Learning agile skills is one of the main challenges in robotics. To this end, reinforcement learning approaches have achieved impressive results. These methods require explicit task information in terms of a reward function or an expert that can be queried in simulation to provide a target control o…

Cited by 73SourceScholar
2022

Learning with Muscles: Benefits for Data-Efficiency and Robustness in Anthropomorphic Tasks

CoRL 2022poster

Humans are able to outperform robots in terms of robustness, versatility, and learning of new tasks in a wide variety of movements. We hypothesize that highly nonlinear muscle dynamics play a large role in providing inherent stability, which is favorable to learning. While recent advances have been…

Cited by 13SourceScholar
2022

On the Pitfalls of Heteroscedastic Uncertainty Estimation with Probabilistic Neural Networks

ICLR 2022poster

Capturing aleatoric uncertainty is a critical part of many machine learning systems. In deep learning, a common approach to this end is to train a neural network to estimate the parameters of a heteroscedastic Gaussian distribution by maximizing the logarithm of the likelihood function under the obs…

2021

Causal Influence Detection for Improving Efficiency in Reinforcement Learning

NeurIPS 2021poster

Many reinforcement learning (RL) environments consist of independent entities that interact sparsely. In such environments, RL agents have only limited influence over other entities in any particular situation. Our idea in this work is that learning can be efficiently guided by knowing when and what…

2021

CombOptNet: Fit the Right NP-Hard Problem by Learning Integer Programming Constraints

ICML 2021spotlight

Bridging logical and algorithmic reasoning with modern machine learning techniques is a fundamental challenge with potentially transformative impact. On the algorithmic side, many NP-hard problems can be expressed as integer programs, in which the constraints play the role of their ’combinatorial sp…

2021

Demystifying Inductive Biases for (Beta-)VAE Based Architectures

ICML 2021spotlight

The performance of Beta-Variational-Autoencoders and their variants on learning semantically meaningful, disentangled representations is unparalleled. On the other hand, there are theoretical arguments suggesting the impossibility of unsupervised disentanglement. In this work, we shed light on the i…

2021

Extracting Strong Policies for Robotics Tasks from Zero-Order Trajectory Optimizers

ICLR 2021poster

Solving high-dimensional, continuous robotic tasks is a challenging optimization problem. Model-based methods that rely on zero-order optimizers like the cross-entropy method (CEM) have so far shown strong performance and are considered state-of-the-art in the model-based reinforcement learning comm…

Cited by 13SourcePDFScholar
2021

Hierarchical Reinforcement Learning with Timed Subgoals

NeurIPS 2021poster

Hierarchical reinforcement learning (HRL) holds great potential for sample-efficient learning on challenging long-horizon tasks. In particular, letting a higher level assign subgoals to a lower level has been shown to enable fast learning on difficult problems. However, such subgoal-based methods ha…

2021

Neuro-algorithmic Policies Enable Fast Combinatorial Generalization

ICML 2021spotlight

Although model-based and model-free approaches to learning the control of systems have achieved impressive results on standard benchmarks, generalization to task variations is still lacking. Recent results suggest that generalization for standard architectures improves only after obtaining exhaustiv…

Cited by 19SourcePDFScholar
2021

Planning from Pixels in Environments with Combinatorially Hard Search Spaces

NeurIPS 2021poster

The ability to form complex plans based on raw visual input is a litmus test for current capabilities of artificial intelligence, as it requires a seamless combination of visual processing and abstract algorithmic execution, two traditionally separate areas of computer science. A recent surge of int…

Cited by 8SourcePDFScholar
2021

Self-supervised Reinforcement Learning with Independently Controllable Subgoals

CoRL 2021poster

To successfully tackle challenging manipulation tasks, autonomous agents must learn a diverse set of skills and how to combine them. Recently, self-supervised agents that set their own abstract goals by exploiting the discovered structure in the environment were shown to perform well on many differe…

Cited by 27SourceScholar
2021

Self-supervised Visual Reinforcement Learning with Object-centric Representations

ICLR 2021spotlight

Autonomous agents need large repertoires of skills to act reasonably on new tasks that they have not seen before. However, acquiring these skills using only a stream of high-dimensional, unstructured, and unlabeled observations is a tricky challenge for any autonomous agent. Previous methods have us…

2021

Sparsely Changing Latent States for Prediction and Planning in Partially Observable Domains

NeurIPS 2021poster

A common approach to prediction and planning in partially observable domains is to use recurrent neural networks (RNNs), which ideally develop and maintain a latent memory about hidden, task-relevant factors. We hypothesize that many of these hidden factors in the physical world are constant over ti…

2020

A Real-Robot Dataset for Assessing Transferability of Learned Dynamics Models

ICRA 2020poster

In the context of model-based reinforcement learning and control, a large number of methods for learning system dynamics have been proposed in recent years. The purpose of these learned models is to synthesize new control policies. An important open question is how robust current dynamics-learning m…

Cited by 10SourceScholar
2020

Deep Graph Matching via Blackbox Differentiation of Combinatorial Solvers

ECCV 2020poster

Building on recent progress at the intersection of combinatorial optimization and deep learning, we propose an end-to-end trainable architecture for deep graph matching that contains unmodified combinatorial solvers. Using the presence of heavily optimized combinatorial solvers together with some im…

2020

Differentiation of Blackbox Combinatorial Solvers

ICLR 2020spotlight

Achieving fusion of deep learning with combinatorial algorithms promises transformative changes to artificial intelligence. One possible approach is to introduce combinatorial building blocks into neural networks. Such end-to-end architectures have the potential to tackle combinatorial problems on r…

Cited by 171SourcecodeScholar
2020

Optimizing Rank-Based Metrics With Blackbox Differentiation

CVPR 2020oral

Rank-based metrics are some of the most widely used criteria for performance evaluation of computer vision models. Despite years of effort, direct optimization for these metrics remains a challenge due to their non-differentiable and non-decomposable nature. We present an efficient, theoretically so…

Cited by 121PDFcodeScholar
2020

Sample-efficient Cross-Entropy Method for Real-time Planning

CoRL 2020

Trajectory optimizers for model-based reinforcement learning, such as the Cross-Entropy Method (CEM), can yield compelling results even in high-dimensional control tasks and sparse-reward environments. However, their sampling inefficiency prevents them from being used for real-time planning and cont

2019

Control What You Can: Intrinsically Motivated Task-Planning Agent

NeurIPS 2019poster

We present a novel intrinsically motivated agent that learns how to control the environment in a sample efficient manner, that is with as few environment interactions as possible, by optimizing learning progress. It learns what can be controlled, how to allocate time and attention as well as the rel…

2016

Compliant control for soft robots: Emergent behavior of a tendon driven anthropomorphic arm

IROS 2016poster

With the accelerated development of robot technologies, optimal control becomes one of the central themes of research. In traditional approaches, the controller, by its internal functionality, finds appropriate actions on the basis of the history of sensor values, guided by the goals, intentions, ob…

Cited by 26SourceScholar