← Search

Michael Bloesch

23 accepted papers

2025

Learning from negative feedback, or positive feedback or both

ICLR 2025spotlight

Existing preference optimization methods often assume scenarios where paired preference feedback (preferred/positive vs. dis-preferred/negative examples) is available. This requirement limits their applicability in scenarios where only unpaired feedback—for example, either positive or negative— is a…

Cited by 0SourcePDFScholar
2024

Imitating Language via Scalable Inverse Reinforcement Learning

NeurIPS 2024poster

The majority of language model training builds on imitation learning. It covers pretraining, supervised fine-tuning, and affects the starting conditions for reinforcement learning from human feedback (RLHF). The simplicity and scalability of maximum likelihood estimation (MLE) for next token predict…

Cited by 8SourcePDFScholar
2024

Mastering Stacking of Diverse Shapes with Large-Scale Iterative Reinforcement Learning on Real Robots

ICRA 2024poster

Reinforcement learning solely from an agent’s self-generated data is often believed to be infeasible for learning on real robots, due to the amount of data needed. However, if done right, agents learning from real data can be surprisingly efficient through re-using previously collected sub-optimal d…

Cited by 6SourceScholar
2024

Offline Actor-Critic Reinforcement Learning Scales to Large Models

ICML 2024oral

We show that offline actor-critic reinforcement learning can scale to large models - such as transformers - and follows similar scaling laws as supervised learning. We find that offline actor-critic algorithms can outperform strong, supervised, behavioral cloning baselines for multi-task training on…

Cited by 17SourcePDFScholar
2021

Towards Real Robot Learning in the Wild: A Case Study in Bipedal Locomotion

CoRL 2021poster

Algorithms for self-learning systems have made considerable progress in recent years, yet safety concerns and the need for additional instrumentation have so far largely limited learning experiments with real robots to well controlled lab settings. In this paper, we demonstrate how a small bipedal r…

Cited by 24SourceScholar
2020

Comparing View-Based and Map-Based Semantic Labelling in Real-Time SLAM

ICRA 2020poster

Generally capable Spatial AI systems must build persistent scene representations where geometric models are combined with meaningful semantic labels. The many approaches to labelling scenes can be divided into two clear groups: view-based which estimate labels from the input view-wise data and then…

Cited by 6SourceScholar
2020

Towards General and Autonomous Learning of Core Skills: A Case Study in Locomotion

CoRL 2020

Modern Reinforcement Learning (RL) algorithms promise to solve difficult motor control problems directly from raw sensory inputs. Their attraction is due in part to the fact that they can represent a general class of methods that allow to learn a solution with a reasonably set reward and minimal pri

Cited by 0SourcePDFScholar
2019

KO-Fusion: Dense Visual SLAM with Tightly-Coupled Kinematic and Odometric Tracking

ICRA 2019poster

Dense visual SLAM methods are able to estimate the 3D structure of an environment and locate the observer within them. They estimate the motion of a camera by matching visual information between consecutive frames, and are thus prone to failure under extreme motion conditions or when observing textu…

Cited by 19SourceScholar
2019

Learning Meshes for Dense Visual SLAM

ICCV 2019poster

Estimating motion and surrounding geometry of a moving camera remains a challenging inference problem. From an information theoretic point of view, estimates should get better as more information is included, such as is done in dense SLAM, but this is strongly dependent on the validity of the underl…

Cited by 28PDFScholar
2019

MID-Fusion: Octree-based Object-Level Multi-Instance Dynamic SLAM

ICRA 2019poster

We propose a new multi-instance dynamic RGB-D SLAM system using an object-level octree-based volumetric representation. It can provide robust camera tracking in dynamic environments and at the same time, continuously estimate geometric, semantic, and motion properties for arbitrary objects in the sc…

Cited by 240SourcecodeScholar
2019

SceneCode: Monocular Dense Semantic Reconstruction Using Learned Encoded Scene Representations

CVPR 2019poster

Systems which incrementally create 3D semantic maps from image sequences must store and update representations of both geometry and semantic entities. However, while there has been much work on the correct formulation for geometrical estimation, state-of-the-art systems usually rely on simple semant…

Cited by 95PDFScholar
2018

CodeSLAM — Learning a Compact, Optimisable Representation for Dense Visual SLAM

CVPR 2018poster

The representation of geometry in real-time 3D perception systems continues to be a critical research issue. Dense maps capture complete surface shape and can be augmented with semantic labels, but their high dimensionality makes them computationally costly to store and process, and unsuitable for r…

Cited by 463SourcePDFScholar
2018

Learning to Solve Nonlinear Least Squares for Monocular Stereo

ECCV 2018poster

Sum-of-squares objective functions are very popular in computer vision algorithms. However, these objective functions are not always easy to optimize. The underlying assumptions made by solvers are often not satisfied and many problems are inherently ill-posed. In this paper, we propose a neural non…

Cited by 100SourcePDFScholar
2018

The Two-State Implicit Filter Recursive Estimation for Mobile Robots

RA-L 2018

This letter deals with recursive filtering for dynamic systems where an explicit process model is not easily devisable. Most Bayesian filters assume the availability of such an explicit process model, and thus may require additional assumptions or fail to properly leverage all available information.

Cited by 60SourceScholar
2016

ANYmal - a highly mobile and dynamic quadrupedal robot

IROS 2016poster

This paper introduces ANYmal, a quadrupedal robot that features outstanding mobility and dynamic motion capability. Thanks to novel, compliant joint modules with integrated electronics, the 30 kg, 0.5 m tall robotic dog is torque controllable and very robust against impulsive loads during running or…

Cited by 1094SourceScholar
2016

Collaborative navigation for flying and walking robots

IROS 2016poster

Flying and walking robots can use their complementary features in terms of viewpoint and payload capability to the best in a heterogeneous team. To this end, we present our online collaborative navigation framework for unknown and challenging terrain. The method leverages the flying robot's onboard…

Cited by 62SourceScholar
2016

Generalized information filtering for MAV parameter estimation

IROS 2016poster

In this paper we present a new estimation algorithm that allows for the combination of information from any number of process and measurement models. This adds more flexibility to the design of the estimator and in our case avoids the need for state augmentation. We achieve this by adapting the maxi…

Cited by 7SourceScholar
2015

Dynamic trotting on slopes for quadrupedal robots

IROS 2015poster

Quadrupedal locomotion on sloped terrains poses different challenges than walking in a mostly flat environment. The robot's configuration needs to be explicitly controlled in order to avoid slipping and kinematic limits. To this end, information about the terrain's inclination is required for carefu…

Cited by 86SourceScholar
2015

Robust visual inertial odometry using a direct EKF-based approach

IROS 2015poster

In this paper, we present a monocular visual-inertial odometry algorithm which, by directly using pixel intensity errors of image patches, achieves accurate tracking performance while exhibiting a very high level of robustness. After detection, the tracking of the multilevel patch features is closel…

Cited by 1155SourceScholar