← Search

Marco Cristani

15 accepted papers

2026

StructXLIP: Enhancing Vision-language Models with Multimodal Structural Cues

CVPR 2026

Edge-based representations are fundamental cues for visual understanding, a principle rooted in early vision research and still central today. We extend this principle to vision-language alignment, showing that isolating and aligning structural cues across modalities can greatly benefit fine-tuning

Cited by 0SourcecodeScholar
2025

Collaborative Instance Object Navigation: Leveraging Uncertainty-Awareness to Minimize Human-Agent Dialogues

ICCV 2025poster

Language-driven instance object navigation assumes that a human initiates the task by providing a detailed description of the target to the embodied agent. While this description is crucial for distinguishing the target from other visually similar instances, providing it prior to navigation can be d…

Cited by 0SourcePDFScholar
2025

LOTS of Fashion! Multi-Conditioning for Image Generation via Sketch-Text Pairing

ICCV 2025poster

Fashion design is a complex creative process that blends visual and textual expressions. Designers convey ideas through sketches, which define spatial structure and design elements, and textual descriptions, capturing material, texture, and stylistic details. In this paper, we present LOcalized Text…

Cited by 0SourcePDFScholar
2025

Seeing the Abstract: Translating the Abstract Language for Vision Language Models

CVPR 2025poster

Natural language goes beyond dryly describing visual content. It contains rich abstract concepts to express feeling, creativity and properties that cannot be directly perceived. Yet, current research in Vision Language Models (VLMs) has not shed light on abstract-oriented language.Our research break…

2025

Towards Real Unsupervised Anomaly Detection Via Confident Meta-Learning

ICCV 2025poster

So-called unsupervised anomaly detection is better described as semi-supervised, as it assumes all training data are nominal. This assumption simplifies training but requires manual data curation, introducing bias and limiting adaptability. We propose Confident Meta-learning (CoMet), a novel trainin…

Cited by 0SourcePDFScholar
2024

Exploring 3D Human Pose Estimation and Forecasting from the Robot’s Perspective: The HARPER Dataset

IROS 2024poster

We introduce HARPER, a novel dataset for 3D body pose estimation and forecasting in dyadic interactions between users and Spot, the quadruped robot manufactured by Boston Dynamics. The key-novelty of HARPER is its focus on the robot’s perspective, i.e., on the data captured by the robot’s sensors. T…

Cited by 3SourceScholar
2024

Mind the Error! Detection and Localization of Instruction Errors in Vision-and-Language Navigation

IROS 2024

Vision-and-Language Navigation in Continuous Environments (VLN-CE) is one of the most intuitive yet challenging embodied AI tasks. Agents are tasked to navigate towards a target goal by executing a set of low-level actions, following a series of natural language instructions. All VLN-CE methods in t

Cited by 13SourceScholar
2022

POP: Mining POtential Performance of New Fashion Products via Webly Cross-Modal Query Expansion

ECCV 2022poster

"We propose a data-centric pipeline able to generate exogenous observation data for the New Fashion Product Performance Forecasting (NFPPF) problem, i.e., predicting the performance of a brand-new clothing probe with no available past observations. Our pipeline manufactures the missing past starting…

2022

Pose Forecasting in Industrial Human-Robot Collaboration

ECCV 2022poster

"Pushing back the frontiers of collaborative robots in industrial environments, we propose a new Separable-Sparse Graph Convolutional Network (SeS-GCN) for pose forecasting. For the first time, SeS-GCN bottlenecks the interaction of the spatial, temporal and channel-wise dimensions in GCNs, and it l…

2022

Spatial Commonsense Graph for Object Localisation in Partial Scenes

CVPR 2022poster

We solve object localisation in partial scenes, a new problem of estimating the unknown position of an object (e.g. where is the bag?) given a partial 3D scan of a scene. The proposed solution is based on a novel scene graph model, the Spatial Commonsense Graph (SCG), where objects are the nodes and…

Cited by 23PDFcodeScholar
2021

POMP++: Pomcp-based Active Visual Search in unknown indoor environments

IROS 2021poster

In this paper, we focus on the problem of learning online an optimal policy for Active Visual Search (AVS) of objects in unknown indoor environments. We propose POMP++, a planning strategy that introduces a novel formulation on top of the classic Partially Observable Monte Carlo Planning (POMCP) fra…

Cited by 17SourceScholar
2020

Leveraging Acoustic Images for Effective Self-Supervised Audio Representation Learning

ECCV 2020poster

In this paper, we propose the use of a new modality characterized by a richer information content, namely acoustic images, for the sake of audio-visual scene understanding. Each pixel in such images is characterized by a spectral signature, associated to a specific direction in space and obtained by…

2018

MX-LSTM: Mixing Tracklets and Vislets to Jointly Forecast Trajectories and Head Poses

CVPR 2018poster

Recent approaches on trajectory forecasting use tracklets to predict the future positions of pedestrians exploiting Long Short Term Memory (LSTM) architectures. This paper shows that adding vislets, that is, short sequences of head pose estimations, allows to increase significantly the trajectory fo…

Cited by 153SourcePDFScholar
2015

The S-Hock Dataset: Analyzing Crowds at the Stadium

CVPR 2015poster

The topic of crowd modeling in computer vision usually assumes a single generic typology of crowd, which is very simplistic. In this paper we adopt a taxonomy that is widely accepted in sociology, focusing on a particular category, the spectator crowd, which is formed by people "interested in watchi…

Cited by 57SourcePDFScholar