← Search

Alexander Mathis

12 accepted papers

2026

KINESIS: Motion Imitation for Human Musculoskeletal Locomotion

ICRA 2026poster

How do humans move? Advances in reinforcement learning (RL) have produced impressive results in capturing human motion using physics-based humanoid control. However, torque-controlled humanoids fail to model key aspects of human motor control such as biomechanical joint constraints & non-linear and …

2026

LLaVAction: evaluating and training multi-modal large language models for action understanding

ICLR 2026poster

Understanding human behavior requires measuring behavioral actions. Due to its complexity, behavior is best mapped onto a rich, semantic structure such as language. Emerging multimodal large language models (MLLMs) are promising candidates, but their fine-grained action understanding ability has not…

Cited by 0SourcecodeScholar
2025

EPFL-Smart-Kitchen: An Ego-Exo Multi-Modal Dataset for Challenging Action and Motion Understanding in Video-Language Models

NeurIPS 2025poster

Understanding behavior requires datasets that capture humans while carrying out complex tasks. The kitchen is an excellent environment for assessing human motor and cognitive function, as many complex actions are naturally exhibited in kitchens from chopping to cleaning. Here, we introduce the EPFL-…

Cited by 0SourcecodeScholar
2025

MammAlps: A Multi-view Video Behavior Monitoring Dataset of Wild Mammals in the Swiss Alps

CVPR 2025highlight

Monitoring wildlife is essential for ecology and ethology, especially in light of the increasing human impact on ecosystems. Camera traps have emerged as habitat-centric sensors enabling the study of wildlife populations at scale with minimal disturbance. However, the lack of annotated video dataset…

2025

PICLe: Pseudo-annotations for In-Context Learning in Low-Resource Named Entity Detection

NAACL 2025long

In-context learning (ICL) enables Large Language Models (LLMs) to perform tasks using few demonstrations, facilitating task adaptation when labeled examples are hard to come by. However, ICL is sensitive to the choice of demonstrations, and it remains unclear which demonstration attributes enable in…

2024

HOISDF: Constraining 3D Hand-Object Pose Estimation with Global Signed Distance Fields

CVPR 2024poster

Human hands are highly articulated and versatile at handling objects. Jointly estimating the 3D poses of a hand and the object it manipulates from a monocular camera is challenging due to frequent occlusions. Thus existing methods often rely on intermediate 3D shape representations to increase perfo…

2024

ODEFormer: Symbolic Regression of Dynamical Systems with Transformers

ICLR 2024spotlight

We introduce ODEFormer, the first transformer able to infer multidimensional ordinary differential equation (ODE) systems in symbolic form from the observation of a single solution trajectory. We perform extensive evaluations on two datasets: (i) the existing ‘Strogatz’ dataset featuring two-dimensi…

2023

AmadeusGPT: a natural language interface for interactive animal behavioral analysis

NeurIPS 2023poster

The process of quantifying and analyzing animal behavior involves translating the naturally occurring descriptive language of their actions into machine-readable code. Yet, codifying behavior analysis is often challenging without deep understanding of animal behavior and technical machine learning k…

2023

Latent exploration for Reinforcement Learning

NeurIPS 2023poster

In Reinforcement Learning, agents learn policies by exploring and interacting with the environment. Due to the curse of dimensionality, learning policies that map high-dimensional sensory input to motor output is particularly challenging. During training, state of the art methods (SAC, PPO, etc.) ex…

2023

Rethinking Pose Estimation in Crowds: Overcoming the Detection Information Bottleneck and Ambiguity

ICCV 2023poster

Frequent interactions between individuals are a fundamental challenge for pose estimation algorithms. Current pipelines either use an object detector together with a pose estimator (top-down approach), or localize all body parts first and then link them to predict the pose of individuals (bottom-up)…

Cited by 28PDFcodeScholar
2022

DMAP: a Distributed Morphological Attention Policy for learning to locomote with a changing body

NeurIPS 2022accept

Biological and artificial agents need to deal with constant changes in the real world. We study this problem in four classical continuous control environments, augmented with morphological perturbations. Learning to locomote when the length and the thickness of different body parts vary is challengi…

2021

AcinoSet: A 3D Pose Estimation Dataset and Baseline Models for Cheetahs in the Wild

ICRA 2021poster

Animals are capable of extreme agility, yet understanding their complex dynamics, which have ecological, biomechanical and evolutionary implications, remains challenging. Being able to study this incredible agility will be critical for the development of next-generation autonomous legged robots. In…

Cited by 66SourcecodeScholar