← Search

Alexandre Alahi

73 accepted papers

2026

Deformable Gaussian Occupancy: Decoupling Rigid and Nonrigid Motion with Factorized Distillation

CVPR 2026

Understanding dynamic 3D environments is essential for safe autonomous driving, particularly when reasoning about human-centric, nonrigid agents. However, existing weakly supervised occupancy prediction frameworks predominantly assume rigid-body motion and rely on simple frame-to-frame offsets, limi

Cited by 0SourcecodeScholar
2026

Fusing Satellite Imagery and Planimetric Maps for Cross-View Localization

ICRA 2026poster

Current cross-view localization methods predominantly rely on satellite imagery as the aerial modality. Although recent work explores planimetric maps (e.g., OpenStreetMap tiles), these approaches often lag in performance. Yet both modalities are widely available and possess complementary properties…

2026

HHI-Assist: A Dataset and Benchmark of Human-Human Interaction in Physical Assistance Scenario

ICRA 2026poster

The increasing labor shortage and aging population underline the need for assistive robots to support human care recipients. To enable safe and responsive assistance, robots require accurate human motion prediction in physical interaction scenarios. However, this remains a challenging task due to th…

2026

JointDiff: Bridging Continuous and Discrete in Multi-Agent Trajectory Generation

ICLR 2026poster

Generative models often treat continuous data and discrete events as separate processes, creating a gap in modeling complex systems where they interact synchronously. To bridge this gap, we introduce $\textbf{JointDiff}$, a novel diffusion framework designed to unify these two processes by simultane…

Cited by 0SourcecodeScholar
2026

LayerSync: Self-aligning Intermediate Layers

ICLR 2026poster

We propose LayerSync, a domain-agnostic approach for improving the generation quality and the training efficiency of diffusion models. Prior studies have highlighted the connection between the quality of generation and the representations learned by diffusion models, showing that external guidance o…

Cited by 0SourcecodeScholar
2026

Loc$^{2}$: Interpretable Cross-View Localization via Depth-Lifted Local Feature Matching

ICLR 2026poster

We propose an accurate and interpretable fine-grained cross-view localization method that estimates the 3 Degrees of Freedom (DoF) pose of a ground-level image by matching its local features with a reference aerial image. Unlike prior approaches that rely on global descriptors or bird’s-eye-view (BE…

Cited by 0SourceScholar
2026

MAD: Motion Appearance Decoupling for efficient Driving World Models

CVPR 2026

Recent video diffusion models generate photorealistic, temporally coherent videos, yet they fall short as reliable world models for autonomous driving, where structured motion and physically consistent interactions are essential. Adapting these generalist video models to driving domains has shown pr

Cited by 0SourcecodeScholar
2026

OSMO: Open-vocabulary Self-eMOtion Tracking

CVPR 2026

We introduce the novel task of egocentric self-emotion tracking, which aims to infer an individual's evolving emotions from egocentric multimodal streams such as voice, visual surroundings, semantic subtext, and eye-tracking signals. To establish this research direction, we present: (1) OSMO dataset

Cited by 0SourcecodeScholar
2026

RAP: 3D Rasterization Augmented End-to-End Planning

ICLR 2026poster

Imitation learning for end-to-end driving trains policies only on expert demonstrations. Once deployed in a closed loop, such policies lack recovery data: small mistakes cannot be corrected and quickly compound into failures. A promising direction is to generate alternative viewpoints and trajectori…

Cited by 0SourcecodeScholar
2026

Rethinking Visual Intelligence: Insights from Video Pretraining

ICML 2026poster

Large language models (LLMs) have demonstrated that large-scale pretraining enables systems to adapt rapidly to new problems with little supervision in the language domain. This success, however, has not translated as effectively to the visual domain, where models, including LLMs, continue to strugg…

Cited by 4SourceScholar
2026

Stable Video Infinity: Infinite-Length Video Generation with Error Recycling

ICLR 2026oral

We propose **Stable Video Infinity (SVI)** that can generate non-looping, ultra-long videos with stable visual quality, while supporting per-clip prompt control and multi-modal conditioning. While existing long-video methods attempt to _**mitigate accumulated errors**_ via handcrafted anti-drifting…

Cited by 0SourcecodeScholar
2025

Boosting Omnidirectional Stereo Matching with a Pre-trained Depth Foundation Model

IROS 2025

Omnidirectional depth perception is essential for mobile robotics applications that require scene understanding across a full 360° field of view. Camera-based setups offer a cost-effective option by using stereo depth estimation to generate dense, high-resolution depth maps without relying on expens

Cited by 0SourcecodeScholar
2025

COARSE: Collaborative Pseudo-Labeling with Coarse Real Labels for Off-Road Semantic Segmentation

IROS 2025

Autonomous off-road navigation faces challenges due to diverse, unstructured environments, requiring robust perception with both geometric and semantic understanding. However, scarce densely labeled semantic data limits generalization across domains. Simulated data helps, but introduces domain adapt

Cited by 0SourceScholar
2025

Certified Human Trajectory Prediction

CVPR 2025poster

Predicting human trajectories is essential for the safe operation of autonomous vehicles, yet current data-driven models often lack robustness in case of noisy inputs such as adversarial examples or imperfect observations. Although some trajectory prediction methods have been developed to provide em…

2025

EvoLM: In Search of Lost Language Model Training Dynamics

NeurIPS 2025oral

Modern language model (LM) training has been divided into multiple stages, making it difficult for downstream developers to evaluate the impact of design choices made at each stage. We present EvoLM, a model suite that enables systematic and transparent analysis of LMs' training dynamics across pre-…

Cited by 0SourceScholar
2025

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control

CVPR 2025poster

We present GEM, a Generalizable Ego-vision Multimodal world model that predicts future frames using a reference frame, sparse features, human poses, and ego-trajectories. Hence, our model has precise control over object dynamics, ego-agent motion and human poses. GEM generates paired RGB and depth o…

2025

GeoDistill: Geometry-Guided Self-Distillation for Weakly Supervised Cross-View Localization

ICCV 2025poster

Cross-view localization, the task of estimating a camera's 3-degrees-of-freedom (3-DoF) pose by aligning ground-level images with aerial images, is crucial for large-scale outdoor applications like autonomous navigation and augmented reality. Existing methods often rely on fully supervised learning,…

2025

HELVIPAD: A Real-World Dataset for Omnidirectional Stereo Depth Estimation

CVPR 2025highlight

Despite progress in stereo depth estimation, omnidirectional imaging remains underexplored, mainly due to the lack of appropriate data. We introduce Helvipad, a real-world dataset for omnidirectional stereo depth estimation, featuring 40K video frames from video sequences across diverse environments…

2025

HHI-Assist: A Dataset and Benchmark of Human-Human Interaction in Physical Assistance Scenario

RA-L 2025

The increasing labor shortage and aging population underline the need for assistive robots to support human care recipients. To enable safe and responsive assistance, robots require accurate human motion prediction in physical interaction scenarios. However, this remains a challenging task due to th

Cited by 1SourceScholar
2025

MotionMap: Representing Multimodality in Human Pose Forecasting

CVPR 2025poster

Human pose forecasting is inherently multimodal since multiple future motions exist for an observed pose sequence. However, learning this multimodality is challenging since the task is ill-posed. To address this issue, we propose an alternative paradigm to make the task well-posed. Additionally, whi…

2025

OSKAR: Omnimodal Self-supervised Knowledge Abstraction and Representation

NeurIPS 2025poster

We present OSKAR, the first multimodal foundation model based on bootstrapped latent feature prediction. Unlike generative or contrastive methods, it avoids memorizing unnecessary details (e.g., pixels), and does not require negative pairs, large memory banks, or hand-crafted augmentations. We propo…

Cited by 0SourcecodeScholar
2025

Probabilistic Collision Risk Estimation for Pedestrian Navigation

IROS 2025

Intelligent devices for supporting persons with vision impairment are becoming more widespread, but they are lacking behind the advancements in intelligent driver assistant system. To make a first step forward, this work discusses the integration of the risk model technology, previously used in auto

Cited by 0SourceScholar
2025

Sim-to-Real Causal Transfer: A Metric Learning Approach to Causally-Aware Interaction Representations

CVPR 2025poster

Modeling spatial-temporal interactions among neighboring agents is at the heart of multi-agent problems such as motion forecasting and crowd navigation. Despite notable progress, it remains unclear to which extent modern representations can capture the causal relationships behind agent interactions.…

2025

Towards Generalizable Trajectory Prediction using Dual-Level Representation Learning and Adaptive Prompting

CVPR 2025poster

Existing vehicle trajectory prediction models struggle with generalizability, prediction uncertainties, and handling complex interactions. It is often due to limitations like complex architectures customized for a specific dataset and inefficient multimodal handling. We propose Perceiver with Regist…

Cited by 1SourcePDFScholar
2025

Towards Self-Supervised Covariance Estimation in Deep Heteroscedastic Regression

ICLR 2025poster

Deep heteroscedastic regression models the mean and covariance of the target distribution through neural networks. The challenge arises from heteroscedasticity, which implies that the covariance is sample dependent and is often unknown. Consequently, recent methods learn the covariance through unsup…

Cited by 0SourcePDFScholar
2025

Unified Human Localization and Trajectory Prediction with Monocular Vision

ICRA 2025

Conventional human trajectory prediction models rely on clean curated data, requiring specialized equipment or manual labeling, which is often impractical for robotic applications. The existing predictors tend to overfit to clean observation affecting their robustness when used with noisy inputs. In

Cited by 2SourcecodeScholar
2024

MaskCLR: Attention-Guided Contrastive Learning for Robust Action Representation Learning

CVPR 2024poster

Current transformer-based skeletal action recognition models tend to focus on a limited set of joints and low-level motion patterns to predict action classes. This results in significant performance degradation under small skeleton perturbations or changing the pose estimator between training and te…

2024

S-JEPA: A Joint Embedding Predictive Architecture for Skeletal Action Recognition

ECCV 2024poster

"Masked self-reconstruction of joints has been shown to be a promising pretext task for self-supervised skeletal action recognition. However, this task focuses on predicting isolated, potentially noisy, joint coordinates, which results in an inefficient utilization of the model capacity. In this pap…

2024

Social-Transmotion: Promptable Human Trajectory Prediction

ICLR 2024poster

Accurate human trajectory prediction is crucial for applications such as autonomous vehicles, robotics, and surveillance systems. Yet, existing models often fail to fully leverage the non-verbal social cues human subconsciously communicate when navigating the space. To address this, we introduce *So…

2024

TIC-TAC: A Framework For Improved Covariance Estimation In Deep Heteroscedastic Regression

ICML 2024poster

Deep heteroscedastic regression involves jointly optimizing the mean and covariance of the predicted distribution using the negative log-likelihood. However, recent works show that this may result in sub-optimal convergence due to the challenges associated with covariance estimation. While the liter…

2024

Toward Reliable Human Pose Forecasting With Uncertainty

RA-L 2024

Recently, there has been an arms race of pose forecasting methods aimed at solving the spatio-temporal task of predicting a sequence of future 3D poses of a person given a sequence of past observed ones. However, the lack of unified benchmarks and limited uncertainty analysis have hindered progress

Cited by 14SourcecodeScholar
2024

Towards Motion Forecasting with Real-World Perception Inputs: Are End-to-End Approaches Competitive?

ICRA 2024poster

Motion forecasting is crucial in enabling autonomous vehicles to anticipate the future trajectories of surrounding agents. To do so, it requires solving mapping, detection, tracking, and then forecasting problems, in a multi-step pipeline. In this complex system, advances in conventional forecasting…

Cited by 19SourcecodeScholar
2024

UniTraj: A Unified Framework for Scalable Vehicle Trajectory Prediction

ECCV 2024poster

"Vehicle trajectory prediction has increasingly relied on data-driven solutions, but their ability to scale to different data domains and the impact of larger dataset sizes on their generalization remain under-explored. While these questions can be studied by employing multiple datasets, it is chall…

2024

When Your AI Becomes a Target: AI Security Incidents and Best Practices

AAAI 2024technical

In contrast to vast academic efforts to study AI security, few real-world reports of AI security incidents exist. Released incidents prevent a thorough investigation of the attackers' motives, as crucial information about the company and AI application is missing. As a consequence, it often remains…

2023

A generic diffusion-based approach for 3D human pose prediction in the wild

ICRA 2023poster

Predicting 3D human poses in real-world scenarios, also known as human pose forecasting, is inevitably subject to noisy inputs arising from inaccurate 3D pose estimations and occlusions. To address these challenges, we propose a diffusion-based approach that can predict given noisy observations. We…

Cited by 46SourcecodeScholar
2023

Real-Time Localization for Closed-Loop Control of Assistive Furniture

RA-L 2023

For people with limited mobility, navigating in cluttered indoor environment is challenging. In this work, we propose a mobile assistive furniture suite that is designed to ease the life of people with special needs in indoor movement. To enable intelligent coordination of this system, a key compone

Cited by 2SourceScholar
2022

Motion Style Transfer: Modular Low-Rank Adaptation for Deep Motion Forecasting

CoRL 2022poster

Deep motion forecasting models have achieved great success when trained on a massive amount of data. Yet, they often perform poorly when training data is limited. To address this challenge, we propose a transfer learning approach for efficiently adapting pre-trained forecasting models to new domains…

Cited by 21SourcecodeScholar
2022

Towards Robust and Adaptive Motion Forecasting: A Causal Representation Perspective

CVPR 2022poster

Learning behavioral patterns from observational data has been a de-facto approach to motion forecasting. Yet, the current paradigm suffers from two shortcomings: brittle under distribution shifts and inefficient for knowledge transfer. In this work, we propose to address these challenges from a caus…

Cited by 68PDFScholar
2022

Vehicle Trajectory Prediction Works, but Not Everywhere

CVPR 2022poster

Vehicle trajectory prediction is nowadays a fundamental pillar of self-driving cars. Both the industry and research communities have acknowledged the need for such a pillar by providing public benchmarks. While state-of-the-art methods are impressive, i.e., they have no off-road prediction, their ge…

Cited by 73PDFcodeScholar
2021

Guest Editorial: Introduction to the Special Issue on Long-Term Human Motion Prediction

RA-L 2021

The articles in this special section focus on long term human motion prediction. This represents a key ability for advanced autonomous systems, especially if they operate in densely crowded and highly dynamic environments. In those settings understanding and anticipating human movements is fundament

Cited by 2SourceScholar
2021

MonStereo: When Monocular and Stereo Meet at the Tail of 3D Human Localization

ICRA 2021poster

Monocular and stereo visions are cost-effective solutions for 3D human localization in the context of self-driving cars or social robots. However, they are usually developed independently and have their respective strengths and limitations. We propose a novel unified learning framework that leverage…

Cited by 10SourcecodeScholar
2021

Safety-Aware Motion Prediction With Unseen Vehicles for Autonomous Driving

ICCV 2021poster

Motion prediction of vehicles is critical but challenging due to the uncertainties in complex environments and the limited visibility caused by occlusions and limited sensor ranges. In this paper, we study a new task, safety-aware motion prediction with unseen vehicles for autonomous driving. Unlike…

Cited by 32PDFcodeScholar
2021

TTT++: When Does Self-Supervised Test-Time Training Fail or Thrive?

NeurIPS 2021poster

Test-time training (TTT) through self-supervised learning (SSL) is an emerging paradigm to tackle distributional shifts. Despite encouraging results, it remains unclear when this approach thrives or fails. In this work, we first provide an in-depth look at its limitations and show that TTT can possi…

2020

DeepSVG: A Hierarchical Generative Network for Vector Graphics Animation

NeurIPS 2020poster

Scalable Vector Graphics (SVG) are ubiquitous in modern 2D interfaces due to their ability to scale to different resolutions. However, despite the success of deep learning-based models applied to rasterized images, the problem of vector graphics representation learning and generation remains largely…

2019

Convolutional Relational Machine for Group Activity Recognition

CVPR 2019poster

We present an end-to-end deep Convolutional Neural Network called Convolutional Relational Machine (CRM) for recognizing group activities that utilizes the information in spatial relations between individual persons in image or video. It learns to produce an intermediate spatial representation (acti…

Cited by 155PDFScholar
2019

Crowd-Robot Interaction: Crowd-Aware Robot Navigation With Attention-Based Deep Reinforcement Learning

ICRA 2019poster

Mobility in an effective and socially-compliant manner is an essential yet challenging task for robots operating in crowded spaces. Recent works have shown the power of deep reinforcement learning techniques to learn socially cooperative policies. However, their cooperation ability deteriorates as t…

Cited by 715SourceScholar
2019

MonoLoco: Monocular 3D Pedestrian Localization and Uncertainty Estimation

ICCV 2019poster

We tackle the fundamentally ill-posed problem of 3D human localization from monocular RGB images. Driven by the limitation of neural networks outputting point estimates, we address the ambiguity in the task by predicting confidence intervals through a loss function based on the Laplace distribution.…

Cited by 152PDFcodeScholar
2018

CAR-Net: Clairvoyant Attentive Recurrent Network

ECCV 2018poster

We present an interpretable framework for path prediction that leverages dependencies between agents' behaviors and their spatial navigation environment. We exploit two sources of information: the past motion trajectory of the agent of interest and a wide top-view image of the navigation scene. We p…

Cited by 177SourcePDFScholar
2018

Social GAN: Socially Acceptable Trajectories With Generative Adversarial Networks

CVPR 2018poster

Understanding human motion behavior is critical for autonomous moving platforms (like self-driving cars and social robots) if they are to navigate human-centric environments. This is challenging because human motion is inherently multimodal: given a history of human motion paths, there are many soci…

2017

Jointly Learning Energy Expenditures and Activities Using Egocentric Multimodal Signals

CVPR 2017poster

Physiological signals such as heart rate can provide valuable information about an individual's state and activity. However, existing work on computer vision has not yet explored leveraging these signals to enhance egocentric video understanding. In this work, we propose a model for reasoning on mul…

Cited by 88PDFScholar
2017

Social Scene Understanding: End-To-End Multi-Person Action Localization and Collective Activity Recognition

CVPR 2017oral

We present a unified framework for understanding human social behaviors in raw image sequences. Our model jointly detects multiple individuals, infers their social actions, and estimates the collective actions with a single feed-forward pass through a neural network. We propose a single architecture…

Cited by 296PDFScholar
2017

Tracking the Untrackable: Learning to Track Multiple Cues With Long-Term Dependencies

ICCV 2017poster

The majority of existing solutions to the Multi-Target Tracking (MTT) problem do not combine cues over a long period of time in a coherent fashion. In this paper, we present an online method that encodes long-term temporal dependencies across multiple cues. One key challenge of tracking methods is t…

Cited by 721PDFScholar
2017

Unsupervised Learning of Long-Term Motion Dynamics for Videos

CVPR 2017poster

We present an unsupervised representation learning approach that compactly encodes the motion dependencies in videos. Given a pair of images from a video clip, our framework learns to predict the long-term 3D motions. To reduce the complexity of the learning framework, we propose to describe the mot…

Cited by 254PDFScholar
2016

Social LSTM: Human Trajectory Prediction in Crowded Spaces

CVPR 2016spotlight

Humans navigate complex crowded environments based on social conventions: they respect personal space, yielding right-of-way and avoid collisions. In our work, we propose a data-driven approach to learn these human-human interactions for predicting their future trajectories. This is in contrast to t…

Cited by 4143PDFScholar