← Search

Stefan Wermter

34 accepted papers

2026

The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level

ICML 2026poster

Mixture-of-Experts (MoE) architectures have become the dominant choice for scaling Large Language Models (LLMs), activating only a subset of parameters per token. While primarily adopted for computational efficiency, it remains an open question whether their sparsity makes them inherently easier to …

Cited by 0SourceScholar
2025

QT-TDM: Planning With Transformer Dynamics Model and Autoregressive Q-Learning

RA-L 2025

Inspired by the success of the Transformer architecture in natural language processing and computer vision, we investigate the use of Transformers in Reinforcement Learning (RL), specifically in modeling the environment's dynamics using Transformer Dynamics Models (TDMs). We evaluate the capabilitie

Cited by 9SourceScholar
2025

Shaken, Not Stirred: A Novel Dataset for Visual Understanding of Glasses in Human-Robot Bartending Tasks

IROS 2025

Datasets for object detection often do not account for enough variety of glasses, due to their transparent and reflective properties. Specifically, open-vocabulary object detectors, widely used in embodied robotic agents, fail to distinguish subclasses of glasses. This scientific gap poses an issue

Cited by 0SourceScholar
2024

Enhancing Zero-Shot Chain-of-Thought Reasoning in Large Language Models through Logic

COLING 2024main

Recent advancements in large language models have showcased their remarkable generalizability across various domains. However, their reasoning abilities still have significant room for improvement, especially when confronted with scenarios requiring multi-step reasoning. Although large language mode…

2024

Improving Speech Emotion Recognition with Unsupervised Speaking Style Transfer

ICASSP 2024accepted

Humans can effortlessly modify various prosodic attributes, such as the placement of stress and the intensity of sentiment, to convey a specific emotion while maintaining consistent linguistic content. Motivated by this capability, we propose EmoAug, a novel style transfer model designed to enhance…

Cited by 0SourceScholar
2024

Inverse Kinematics for Neuro-Robotic Grasping with Humanoid Embodied Agents

IROS 2024poster

This paper introduces a novel zero-shot motion planning method that allows users to quickly design smooth robot motions in Cartesian space. A Bézier curve-based Cartesian plan is transformed into a joint space trajectory by our neuro-inspired inverse kinematics (IK) method CycleIK, for which we enab…

Cited by 3SourcecodeScholar
2023

Chat with the Environment: Interactive Multimodal Perception Using Large Language Models

IROS 2023poster

Programming robot behavior in a complex world faces challenges on multiple levels, from dextrous low-level skills to high-level planning and reasoning. Recent pre-trained Large Language Models (LLMs) have shown remarkable reasoning ability in few-shot robotic planning. However, it remains challengin…

Cited by 81SourcecodeScholar
2023

Internally Rewarded Reinforcement Learning

ICML 2023poster

We study a class of reinforcement learning problems where the reward signals for policy learning are generated by a discriminator that is dependent on and jointly optimized with the policy. This interdependence between the policy and the discriminator leads to an unstable learning process because re…

2023

Partially Adaptive Multichannel Joint Reduction of Ego-Noise and Environmental Noise

ICASSP 2023accepted

Human-robot interaction relies on a noise-robust audio processing module capable of estimating target speech from audio recordings impacted by environmental noise, as well as self-induced noise, so-called ego-noise. While external ambient noise sources vary from environment to environment, ego-noise…

Cited by 0SourceScholar
2023

Sample-Efficient Real-Time Planning with Curiosity Cross-Entropy Method and Contrastive Learning

IROS 2023poster

Model-based reinforcement learning (MBRL) with real-time planning has shown great potential in locomotion and manipulation control tasks. However, the existing planning methods, such as the Cross-Entropy Method (CEM), do not scale well to complex high-dimensional environments. One of the key reasons…

Cited by 3SourcecodeScholar
2023

Visually Grounded Commonsense Knowledge Acquisition

AAAI 2023technical

Large-scale commonsense knowledge bases empower a broad range of AI applications, where the automatic extraction of commonsense knowledge (CKE) is a fundamental and challenging problem. CKE from text is known for suffering from the inherent sparsity and reporting bias of commonsense in text. Visual…

2023

Visually Grounded Continual Language Learning with Selective Specialization

EMNLP 2023long findings

A desirable trait of an artificial agent acting in the visual world is to continually learn a sequence of language-informed tasks while striking a balance between sufficiently specializing in each task and building a generalized knowledge for transfer. Selective specialization, i.e., a careful selec…

Cited by 0SourcecodeScholar
2022

Impact Makes a Sound and Sound Makes an Impact: Sound Guides Representations and Explorations

IROS 2022poster

Sound is one of the most informative and abundant modalities in the real world while being robust to sense without contacts by small and cheap sensors that can be placed on mobile devices. Although deep learning is capable of extracting information from multiple sensory inputs, there has been little…

Cited by 12SourcecodeScholar
2022

Integrating Statistical Uncertainty into Neural Network-Based Speech Enhancement

ICASSP 2022accepted

Speech enhancement in the time-frequency domain is often performed by estimating a multiplicative mask to extract clean speech. However, most neural network-based methods perform point estimation, i.e., their output consists of a single mask. In this paper, we study the benefits of modeling uncertai…

Cited by 0SourceScholar
2022

What is Right for Me is Not Yet Right for You: A Dataset for Grounding Relative Directions via Multi-Task Learning

IJCAI 2022poster

Understanding spatial relations is essential for intelligent agents to act and communicate in the physical world. Relative directions are spatial relations that describe the relative positions of target objects with regard to the intrinsic orientation of reference objects. Grounding relative directi…

2021

Robotic Occlusion Reasoning for Efficient Object Existence Prediction

IROS 2021poster

Reasoning about potential occlusions is essential for robots to efficiently predict whether an object exists in an environment. Though existing work shows that a robot with active perception can achieve various tasks, it is still unclear if occlusion reasoning can be achieved. To answer this questio…

Cited by 8SourceScholar
2021

Variational Autoencoder for Speech Enhancement with a Noise-Aware Encoder

ICASSP 2021accepted

Recently, a generative variational autoencoder (VAE) has been proposed for speech enhancement to model speech statistics. However, this approach only uses clean speech in the training phase, making the estimation particularly sensitive to noise presence, especially in low signal-to-noise ratios (SNR…

Cited by 0SourceScholar
2021

Visual Distant Supervision for Scene Graph Generation

ICCV 2021poster

Scene graph generation aims to identify objects and their relations in images, providing structured image representations that can facilitate numerous applications in computer vision. However, scene graph models usually require supervised learning on large quantities of labeled data with intensive h…

Cited by 51PDFcodeScholar
2019

A Personalized Affective Memory Model for Improving Emotion Recognition

ICML 2019oral

Recent models of emotion recognition strongly rely on supervised deep learning solutions for the distinction of general emotion expressions. However, they are not reliable when recognizing online and personalized facial expressions, e.g., for person-specific affective understanding. In this paper, w…

Cited by 37SourcePDFScholar
2019

Designing a Personality-Driven Robot for a Human-Robot Interaction Scenario

ICRA 2019poster

In this paper, we present an autonomous AI system designed for a Human-Robot Interaction (HRI) study, set around a dice game scenario. We conduct a case study to answer our research question: Does a robot with a socially engaged personality lead to a higher acceptance than a competitive personality?…

Cited by 18SourceScholar
2019

Exploring Low-level and High-level Transfer Learning for Multi-task Facial Recognition with a Semi-supervised Neural Network

IROS 2019poster

Facial recognition tasks like identity, age, gender, and emotion recognition received substantial attention in recent years. Their deployment in robotic platforms became necessary for the characterization of most of the non-verbal Human-Robot Interaction (HRI) scenarios. In this regard, deep convolu…

Cited by 1SourceScholar
2019

Generating Multiple Objects at Spatially Distinct Locations

ICLR 2019poster

Recent improvements to Generative Adversarial Networks (GANs) have made it possible to generate realistic images in high resolution based on natural language descriptions such as image captions. Furthermore, conditional GANs allow us to control the image generation process through labels or even nat…

2019

Incorporating End-to-End Speech Recognition Models for Sentiment Analysis

ICRA 2019poster

Previous work on emotion recognition demonstrated a synergistic effect of combining several modalities such as auditory, visual, and transcribed text to estimate the affective state of a speaker. Among these, the linguistic modality is crucial for the evaluation of an expressed emotion. However, man…

Cited by 32SourceScholar
2018

A Neurorobotic Experiment for Crossmodal Conflict Resolution in Complex Environments

IROS 2018poster

Crossmodal conflict resolution is crucial for robot sensorimotor coupling through the interaction with the environment, yielding swift and robust behaviour also in noisy conditions. In this paper, we propose a neurorobotic experiment in which an iCub robot exhibits human-like responses in a complex…

Cited by 16SourceScholar
2018

An Ensemble with Shared Representations Based on Convolutional Networks for Continually Learning Facial Expressions

IROS 2018poster

Social robots able to continually learn facial expressions could progressively improve their emotion recognition capability towards people interacting with them. Semi-supervised learning through ensemble predictions is an efficient strategy to leverage the high exposure of unlabelled facial expressi…

Cited by 13SourceScholar
2018

Deep Neural Object Analysis by Interactive Auditory Exploration with a Humanoid Robot

IROS 2018poster

We present a novel approach for interactive auditory object analysis with a humanoid robot. The robot elicits sensory information by physically shaking visually indistinguishable plastic capsules. It gathers the resulting audio signals from microphones that are embedded into the robotic ears. A neur…

Cited by 21SourceScholar
2018

EmoRL: Continuous Acoustic Emotion Classification Using Deep Reinforcement Learning

ICRA 2018poster

Acoustically expressed emotions can make communication with a robot more efficient. Detecting emotions like anger could provide a clue for the robot indicating unsafe/undesired situations. Recently, several deep neural network-based models have been proposed which establish new state-of-the-art resu…

Cited by 31SourceScholar
2018

Hear the Egg - Demonstrating Robotic Interactive Auditory Perception

IROS 2018poster

We present an illustrative example of an interactive auditory perception approach performed by a humanoid robot called NICO, the Neuro Inspired COmpanion [1]. The video demonstrates a material classification task in the style of a classic TV game show. NICO and another candidate are supposed to dete…

Cited by 5SourceScholar
2018

Object Detection and Pose Estimation Based on Convolutional Neural Networks Trained with Synthetic Data

IROS 2018poster

Instance-based object detection and fine pose estimation is an active research problem in computer vision. While the traditional interest-point-based approaches for pose estimation are precise, their applicability in robotic tasks relies on controlled environments and rigid objects with detailed tex…

Cited by 51SourceScholar
2018

On the Robustness of Speech Emotion Recognition for Human-Robot Interaction with Deep Neural Networks

IROS 2018poster

Speech emotion recognition (SER) is an important aspect of effective human-robot collaboration and received a lot of attention from the research community. For example, many neural network-based architectures were proposed recently and pushed the performance to a new level. However, the applicabilit…

Cited by 76SourceScholar
2016

Multi-modal integration of dynamic audiovisual patterns for an interactive reinforcement learning scenario

IROS 2016poster

Robots in domestic environments are receiving more attention, especially in scenarios where they should interact with parent-like trainers for dynamically acquiring and refining knowledge. A prominent paradigm for dynamically learning new tasks has been reinforcement learning. However, due to excess…

Cited by 45SourceScholar
2015

Recognizing complex mental states with deep hierarchical features for Human-Robot Interaction

IROS 2015poster

The use of emotional states for Human-Robot Interaction (HRI) has attracted considerable attention in recent years. One of the most challenging tasks is to recognize the spontaneous expression of emotions, especially in an HRI scenario. Every person has a different way to express emotions, and this…

Cited by 1SourceScholar