← Search

Simon Stepputtis

27 accepted papers

2026

LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models

AAAI 2026technical

Large Language Models (LLMs) exhibit impressive general-purpose capabilities but also introduce serious safety risks, particularly the potential for deception as models acquire increased agency and human oversight diminishes. In this work, we present LieCraft: a novel evaluation framework and sandbo

Cited by 0SourcePDFScholar
2025

Adaptively Coordinating with Novel Partners via Learned Latent Strategies

NeurIPS 2025poster

Adaptation is the cornerstone of effective collaboration among heterogeneous team members. In human-agent teams, artificial agents need to adapt to their human partners in real time, as individuals often have unique preferences and policies that may change dynamically throughout interactions. This b…

Cited by 0SourceScholar
2025

InstructPart: Task-Oriented Part Segmentation with Instruction Reasoning

ACL 2025long

Large multimodal foundation models, particularly in the domains of language and vision, have significantly advanced various tasks, including robotics, autonomous driving, information retrieval, and grounding. However, many of these models perceive objects as indivisible, overlooking the components t…

2025

OMG: Opacity Matters in Material Modeling with Gaussian Splatting

ICLR 2025poster

Decomposing geometry, materials and lighting from a set of images, namely inverse rendering, has been a long-standing problem in computer vision and graphics. Recent advances in neural rendering enable photo-realistic and plausible inverse rendering results. The emergence of 3D Gaussian Splatting ha…

Cited by 0SourcePDFScholar
2025

ONLY: One-Layer Intervention Sufficiently Mitigates Hallucinations in Large Vision-Language Models

ICCV 2025poster

Recent Large Vision-Language Models (LVLMs) have introduced a new paradigm for understanding and reasoning about image input through textual responses. Although they have achieved remarkable performance across a range of multi-modal tasks, they face the persistent challenge of hallucination, which i…

2025

Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models

ICLR 2025poster

While recent Large Vision-Language Models (LVLMs) have shown remarkable performance in multi-modal tasks, they are prone to generating hallucinatory text responses that do not align with the given visual input, which restricts their practical applicability in real-world scenarios. In this work, insp…

2024

A Comparison of Imitation Learning Algorithms for Bimanual Manipulation

RA-L 2024

Amidst the wide popularity of imitation learning algorithms in robotics, their properties regarding hyperparameter sensitivity, ease of training, data efficiency, and performance have not been well-studied in high-precision industry-inspired environments. In this work, we demonstrate the limitations

Cited by 22SourceScholar
2024

Dual Prototype Evolving for Test-Time Generalization of Vision-Language Models

NeurIPS 2024poster

Test-time adaptation, which enables models to generalize to diverse data with unlabeled test samples, holds significant value in real-world scenarios. Recently, researchers have applied this setting to advanced pre-trained vision-language models (VLMs), developing approaches such as test-time prompt…

2024

GL-NeRF: Gauss-Laguerre Quadrature Enables Training-Free NeRF Acceleration

NeurIPS 2024poster

Volume rendering in neural radiance fields is inherently time-consuming due to the large number of MLP calls on the points sampled per ray. Previous works would address this issue by introducing new neural networks or data structures. In this work, we propose GL-NeRF, a new perspective of computing…

2024

HiKER-SGG: Hierarchical Knowledge Enhanced Robust Scene Graph Generation

CVPR 2024poster

Being able to understand visual scenes is a precursor for many downstream tasks including autonomous driving robotics and other vision-based approaches. A common approach enabling the ability to reason over visual data is Scene Graph Generation (SGG); however many existing approaches assume undistur…

2024

Let Me Help You! Neuro-Symbolic Short-Context Action Anticipation

RA-L 2024

In an era where robots become available to the general public, the applicability of assistive robotics extends across numerous aspects of daily life, including in-home robotics. This work presents a novel approach for such systems, leveraging long-horizon action anticipation from short-observation c

Cited by 5SourceScholar
2024

LogiCity: Advancing Neuro-Symbolic AI with Abstract Urban Simulation

NeurIPS 2024poster

Recent years have witnessed the rapid development of Neuro-Symbolic (NeSy) AI systems, which integrate symbolic reasoning into deep neural networks. However, most of the existing benchmarks for NeSy AI fail to provide long-horizon reasoning tasks with complex multi-agent interactions. Furthermore, t…

2024

Navigating Noisy Feedback: Enhancing Reinforcement Learning with Error-Prone Language Models

EMNLP 2024finding

The correct specification of reward models is a well-known challenge in reinforcement learning.Hand-crafted reward functions often lead to inefficient or suboptimal policies and may not be aligned with user values.Reinforcement learning from human feedback is a successful technique that can mitigate…

2024

ShapeGrasp: Zero-Shot Task-Oriented Grasping with Large Language Models through Geometric Decomposition

IROS 2024poster

Task-oriented grasping of unfamiliar objects is a necessary skill for robots in dynamic in-home environments. Inspired by the human capability to grasp such objects through intuition about their shape and structure, we present a novel zero-shot task-oriented grasping method leveraging a geometric de…

Cited by 10SourcecodeScholar
2023

Characterizing Out-of-Distribution Error via Optimal Transport

NeurIPS 2023poster

Out-of-distribution (OOD) data poses serious challenges in deployed machine learning models, so methods of predicting a model's performance on OOD data without labels are important for machine learning safety. While a number of methods have been proposed by prior work, they often underestimate the a…

Cited by 20SourcePDFScholar
2023

Explainable Action Advising for Multi-Agent Reinforcement Learning

ICRA 2023poster

Action advising is a knowledge transfer technique for reinforcement learning based on the teacher-student paradigm. An expert teacher provides advice to a student during training in order to improve the student's sample efficiency and policy performance. Such advice is commonly given in the form of…

Cited by 24SourcecodeScholar
2023

Long-Horizon Dialogue Understanding for Role Identification in the Game of Avalon with Large Language Models

EMNLP 2023long findings

Deception and persuasion play a critical role in long-horizon dialogues between multiple parties, especially when the interests, goals, and motivations of the participants are not aligned. Such complex tasks pose challenges for current Large Language Models (LLM) as deception and persuasion can easi…

Cited by 0SourcecodeScholar
2023

Theory of Mind for Multi-Agent Collaboration via Large Language Models

EMNLP 2023long main

While Large Language Models (LLMs) have demonstrated impressive accomplishments in both reasoning and planning, their abilities in multi-agent collaborations remains largely unexplored. This study evaluates LLM-based agents in a multi-agent cooperative text game with Theory of Mind (ToM) inference t…

Cited by 0SourcecodeScholar
2022

A System for Imitation Learning of Contact-Rich Bimanual Manipulation Policies

IROS 2022poster

In this paper, we discuss a framework for teaching bimanual manipulation tasks by imitation. To this end, we present a system and algorithms for learning compliant and contact-rich robot behavior from human demonstrations. The presented system combines insights from admittance control and machine le…

Cited by 34SourceScholar
2022

Concept Learning for Interpretable Multi-Agent Reinforcement Learning

CoRL 2022poster

Multi-agent robotic systems are increasingly operating in real-world environments in close proximity to humans, yet are largely controlled by policy models with inscrutable deep neural network representations. We introduce a method for incorporating interpretable concepts from a domain expert into m…

Cited by 23SourceScholar
2022

Modularity through Attention: Efficient Training and Transfer of Language-Conditioned Policies for Robot Manipulation

CoRL 2022poster

Language-conditioned policies allow robots to interpret and execute human instructions. Learning such policies requires a substantial investment with regards to time and compute resources. Still, the resulting controllers are highly device-specific and cannot easily be transferred to a robot with di…

Cited by 25SourcecodeScholar
2020

Language-Conditioned Imitation Learning for Robot Manipulation Tasks

NeurIPS 2020spotlight

Imitation learning is a popular approach for teaching motor skills to robots. However, most approaches focus on extracting policy parameters from execution traces alone (i.e., motion trajectories and perceptual data). No adequate communication channel exists between the human expert and the robot to…

2019

Improved Exploration through Latent Trajectory Optimization in Deep Deterministic Policy Gradient

IROS 2019poster

Model-free reinforcement learning algorithms such as Deep Deterministic Policy Gradient (DDPG) often require additional exploration strategies, especially if the actor is of deterministic nature. This work evaluates the use of model-based trajectory optimization methods used for exploration in Deep…

Cited by 15SourceScholar
2019

Learning Interactive Behaviors for Musculoskeletal Robots Using Bayesian Interaction Primitives

IROS 2019poster

Musculoskeletal robots that are based on pneumatic actuation have a variety of properties, such as compliance and back-drivability, that render them particularly appealing for human-robot collaboration. However, programming interactive and responsive behaviors for such systems is extremely challengi…

Cited by 21SourceScholar
2019

Probabilistic Multimodal Modeling for Human-Robot Interaction Tasks

RSS 2019poster

Human-robot interaction benefits greatly from multimodal sensor inputs as they enable increased robustness and generalization accuracy. Despite this observation, few HRI methods are capable of efficiently performing inference for multimodal systems. In this work, we introduce a reformulation of Inte…

Cited by 33SourcePDFScholar
2018

Extrinsic Dexterity Through Active Slip Control Using Deep Predictive Models

ICRA 2018poster

We present a machine learning methodology for actively controlling slip, in order to increase robot dexterity. Leveraging recent insights in deep learning, we propose a Deep Predictive Model that uses tactile sensor information to reason about slip and its future influence on the manipulated object.…

Cited by 13SourceScholar
2017

A system for learning continuous human-robot interactions from human-human demonstrations

ICRA 2017poster

We present a data-driven imitation learning system for learning human-robot interactions from human-human demonstrations. During training, the movements of two interaction partners are recorded through motion capture and an interaction model is learned. At runtime, the interaction model is used to c…

Cited by 106SourceScholar