← Search

Tim Seyde

14 accepted papers

2026

ZIP-RC: Zero-overhead Inference-time Prediction of Reward and Cost for Adaptive and Interpretable Generation

ICLR 2026poster

Large language models excel at reasoning but lack key aspects of introspection, including the ability to anticipate their own success and the computation required to achieve it. Humans use real-time introspection to decide how much effort to invest, when to make multiple attempts, when to stop, and…

Cited by 0SourceScholar
2023

Dynamic Multi-Team Racing: Competitive Driving on 1/10-th Scale Vehicles via Learning in Simulation

CoRL 2023poster

Autonomous racing is a challenging task that requires vehicle handling at the dynamic limits of friction. While single-agent scenarios like Time Trials are solved competitively with classical model-based or model-free feedback control, multi-agent wheel-to-wheel racing poses several challenges inclu…

Cited by 6SourceScholar
2023

Gigastep - One Billion Steps per Second Multi-agent Reinforcement Learning

NeurIPS 2023poster

Multi-agent reinforcement learning (MARL) research is faced with a trade-off: it either uses complex environments requiring large compute resources, which makes it inaccessible to researchers with limited resources, or relies on simpler dynamics for faster execution, which makes the transferability…

2023

Measuring Interpretability of Neural Policies of Robots with Disentangled Representation

CoRL 2023oral

The advancement of robots, particularly those functioning in complex human-centric environments, relies on control solutions that are driven by machine learning. Understanding how learning-based controllers make decisions is crucial since robots are mostly safety-critical systems. This urges a forma…

Cited by 8SourceScholar
2023

Solving Continuous Control via Q-learning

ICLR 2023poster

While there has been substantial success for solving continuous control with actor-critic methods, simpler critic-only methods such as Q-learning find limited application in the associated high-dimensional action spaces. However, most actor-critic methods come at the cost of added complexity: heuris…

2023

Towards Cooperative Flight Control Using Visual-Attention

IROS 2023poster

The cooperation of a human pilot with an autonomous agent during flight control realizes parallel autonomy. We propose an air-guardian system that facilitates cooperation between a pilot with eye tracking and a parallel end-to-end neural control system. Our vision-based air-guardian system combines…

Cited by 7SourceScholar
2022

Interpretable Autonomous Flight Via Compact Visualizable Neural Circuit Policies

RA-L 2022

We learn interpretable end-to-end controllers based on Neural Circuit Policies (NCPs) to enable goal reaching and dynamic obstacle avoidance in flight domains. In addition to being able to learn high-quality control, NCP networks are designed with a small number of neurons. This property allows for

Cited by 8SourceScholar
2021

Is Bang-Bang Control All You Need? Solving Continuous Control with Bernoulli Policies

NeurIPS 2021poster

Reinforcement learning (RL) for continuous control typically employs distributions whose support covers the entire action space. In this work, we investigate the colloquially known phenomenon that trained agents often prefer actions at the boundaries of that space. We draw theoretical connections to…

Cited by 52SourcePDFScholar
2021

Learning to Plan Optimistically: Uncertainty-Guided Deep Exploration via Latent Model Ensembles

CoRL 2021poster

Learning complex robot behaviors through interaction requires structured exploration. Planning should target interactions with the potential to optimize long-term performance, while only reducing uncertainty where conducive to this objective. This paper presents Latent Optimistic Value Exploration…

Cited by 12SourceScholar
2021

Strength Through Diversity: Robust Behavior Learning via Mixture Policies

CoRL 2021poster

Efficiency in robot learning is highly dependent on hyperparameters. Robot morphology and task structure differ widely and finding the optimal setting typically requires sequential or parallel repetition of experiments, strongly increasing the interaction count. We propose a training method that onl…

Cited by 10SourceScholar
2020

Deep Latent Competition: Learning to Race Using Visual Control Policies in Latent Space

CoRL 2020

Learning competitive behaviors in multi-agent settings such as racing requires long-term reasoning about potential adversarial interactions. This paper presents Deep Latent Competition (DLC), a novel reinforcement learning algorithm that learns competitive visual control policies through self-play i

2019

Locomotion Planning through a Hybrid Bayesian Trajectory Optimization

ICRA 2019poster

Locomotion planning for legged systems requires reasoning about suitable contact schedules. The contact sequence and timings constitute a hybrid dynamical system and prescribe a subset of achievable motions. State-of-the-art approaches cast motion planning as an optimal control problem. In order to…

Cited by 18SourceScholar
2018

Inclusion of Angular Momentum During Planning for Capture Point Based Walking

ICRA 2018poster

When walking at high speeds, the swing legs of robots produce a non-negligible angular momentum rate. To accommodate this, we provide a reference trajectory generator for bipedal walking that incorporates predicted centroidal angular momentum at the planning stage. This can be done efficiently as th…

Cited by 30SourceScholar