← Search

Lamberto Ballan

11 accepted papers

2026

PersONAL: Towards a Comprehensive Benchmark for Personalized Embodied Agents

ICRA 2026poster

Recent advances in Embodied AI have enabled agents to perform increasingly complex tasks and adapt to diverse environments. However, deploying such agents in realistic human-centered scenarios, such as domestic households, remains challenging, particularly due to the difficulty of modeling individua…

2025

Following the Human Thread in Social Navigation

ICLR 2025spotlight

The success of collaboration between humans and robots in shared environments relies on the robot's real-time adaptation to human motion. Specifically, in Social Navigation, the agent should be close enough to assist but ready to back up to let the human move freely, avoiding collisions. Human traje…

2025

TANGO: Training-free Embodied AI Agents for Open-world Tasks

CVPR 2025poster

Large Language Models (LLMs) have demonstrated excellent capabilities in composing various modules together to create programs that can perform complex reasoning tasks on images. In this paper, we propose TANGO, an approach that extends the program composition via LLMs already observed for images, a…

Cited by 0SourcePDFScholar
2024

Distilling Knowledge for Short-to-Long Term Trajectory Prediction

IROS 2024

Long-term trajectory forecasting is an important and challenging problem in the fields of computer vision, machine learning, and robotics. One fundamental difficulty stands in the evolution of the trajectory that becomes more and more uncertain and unpredictable as the time horizon grows, subsequent

Cited by 5SourceScholar
2023

Exploiting Proximity-Aware Tasks for Embodied Social Navigation

ICCV 2023poster

Learning how to navigate among humans in an occluded and spatially constrained indoor environment, is a key ability required to embodied agents to be integrated into our society. In this paper, we propose an end-to-end architecture that exploits Proximity-Aware Tasks (referred as to Risk and Proximi…

Cited by 15PDFcodeScholar
2023

TAMformer: Multi-Modal Transformer with Learned Attention Mask for Early Intent Prediction

ICASSP 2023accepted

Human intention prediction is a growing area of research where an activity in a video has to be anticipated by a vision-based system. To this end, the model creates a representation of the past, and subsequently, it produces future hypotheses about upcoming scenarios. In this work, we focus on pedes…

Cited by 0SourceScholar
2022

How Many Observations Are Enough? Knowledge Distillation for Trajectory Forecasting

CVPR 2022poster

Accurate prediction of future human positions is an essential task for modern video-surveillance systems. Current state-of-the-art models usually rely on a "history" of past tracked locations (e.g., 3 to 5 seconds) to predict a plausible sequence of future locations (e.g., up to the next 5 seconds).…

Cited by 78PDFScholar
2022

Online Learning of Reusable Abstract Models for Object Goal Navigation

CVPR 2022poster

In this paper, we present a novel approach to incrementally learn an Abstract Model of an unknown environment, and show how an agent can reuse the learned model for tackling the Object Goal Navigation task. The Abstract Model is a finite state machine in which each state is an abstraction of a state…

Cited by 26PDFScholar
2021

Conditional Variational Capsule Network for Open Set Recognition

ICCV 2021poster

In open set recognition, a classifier has to detect unknown classes that are not known at training time. In order to recognize new categories, the classifier has to project the input samples of known classes in very compact and separated regions of the features space for discriminating samples of un…

Cited by 66PDFcodeScholar
2017

Effective Fisher vector aggregation for 3D object retrieval

ICASSP 2017accepted

We formulate the task of 3D object retrieval as a visual search problem where a database containing videos of objects captured manually from different viewpoints is queried using a single image. We propose to aggregate visual information of similar views and use the Fisher vector (FV) framework to c…

Cited by 0SourceScholar