← Search

Alessio Del Bue

38 accepted papers

2026

Directed Semi-Simplicial Learning with Applications to Brain Activity Decoding

ICLR 2026poster

Graph Neural Networks (GNNs) excel at learning from pairwise interactions but often overlook multi-way and hierarchical relationships. Topological Deep Learning (TDL) addresses this limitation by leveraging combinatorial topological spaces, such as simplicial or cell complexes. However, existing TDL…

Cited by 0SourcecodeScholar
2026

Sheaves Reloaded: A Direction Awakening

ICLR 2026poster

Sheaf Neural Networks (SNNs) are a powerful algebraic-topology generalization of Graph Neural Networks (GNNs), and have been shown to significantly improve our ability to model complex relational data. While the GNN literature proved that incorporating directionality can substantially boost performa…

Cited by 0SourcecodeScholar
2025

BillBoard Splatting (BBSplat): Learnable Textured Primitives for Novel View Synthesis

ICCV 2025poster

We present billboard Splatting (BBSplat) - a novel approach for novel view synthesis based on textured geometric primitives. BBSplat represents the scene as a set of optimizable textured planar primitives with learnable RGB textures and alpha-maps to control their shape. BBSplat primitives can be us…

2025

Embodied Image Captioning: Self-supervised Learning Agents for Spatially Coherent Image Descriptions

ICCV 2025poster

We present a self-supervised method to improve an agent's abilities in describing arbitrary objects while actively exploring a generic environment. This is a challenging problem, as current models struggle to obtain coherent image captions due to different camera viewpoints and clutter. We propose a…

2025

GRASPLAT: Enabling dexterous grasping through novel view synthesis

IROS 2025

Achieving dexterous robotic grasping with multi-fingered hands remains a significant challenge. While existing methods rely on complete 3D scans to predict grasp poses, these approaches face limitations due to the difficulty of acquiring high-quality 3D data in real-world scenarios. In this paper, w

Cited by 0SourcecodeScholar
2025

Measuring Uncertainty in Shape Completion to Improve Grasp Quality

IROS 2025

Shape completion networks have been used recently in real-world robotic experiments to complete the missing/hidden information in environments where objects are only observed in one or few instances where self-occlusions are bound to occur. Nowadays, most approaches rely on deep neural networks that

Cited by 1SourcecodeScholar
2025

Reasoning in Visual Navigation of End-to-end Trained Agents: A Dynamical Systems Approach

CVPR 2025highlight

Progress in Embodied AI has made it possible for end-to-end-trained agents to navigate in photo-realistic environments with high-level reasoning and zero-shot or language-conditioned behavior, but evaluations and benchmarks are still dominated by simulation. In this work, we focus on the fine-graine…

2025

ReassembleNet: Learnable Keypoints and Diffusion for 2D Fresco Reconstruction

ICCV 2025poster

The task of reassembly is a significant challenge across multiple domains, including archaeology, genomics, and molecular docking, requiring the precise placement and orientation of elements to reconstruct an original structure. In this work, we address key limitations in state-of-the-art Deep Learn…

Cited by 0SourcePDFScholar
2024

6DGS: 6D Pose Estimation from a Single Image and a 3D Gaussian Splatting Model

ECCV 2024poster

"We propose to estimate the camera pose of a target RGB image given a 3D Gaussian Splatting (3DGS) model representing the scene. avoids the iterative process typical of analysis-by-synthesis methods (iNeRF) that also require an initialization of the camera pose in order to converge. Instead, our met…

2024

DiffAssemble: A Unified Graph-Diffusion Model for 2D and 3D Reassembly

CVPR 2024poster

Reassembly tasks play a fundamental role in many fields and multiple approaches exist to solve specific reassembly problems. In this context we posit that a general unified model can effectively address them all irrespective of the input data type (image 3D etc.). We introduce DiffAssemble a Graph N…

2024

IFFNeRF: Initialisation Free and Fast 6DoF pose estimation from a single image and a NeRF model

ICRA 2024poster

We introduce IFFNeRF to estimate the six degrees-of-freedom (6DoF) camera pose of a given image, building on the Neural Radiance Fields (NeRF) formulation. IFFNeRF is specifically designed to operate in real-time and eliminates the need for an initial pose guess that is proximate to the sought solut…

Cited by 7SourcecodeScholar
2024

Look Around and Learn: Self-Training Object Detection by Exploration

ECCV 2024poster

"When an object detector is deployed in a novel setting it often experiences a drop in performance. This paper studies how an embodied agent can automatically fine-tune a pre-existing object detector while exploring and acquiring images in a new environment without relying on human intervention, i.e…

2024

Mind the Error! Detection and Localization of Instruction Errors in Vision-and-Language Navigation

IROS 2024

Vision-and-Language Navigation in Continuous Environments (VLN-CE) is one of the most intuitive yet challenging embodied AI tasks. Agents are tasked to navigate towards a target goal by executing a set of low-level actions, following a series of natural language instructions. All VLN-CE methods in t

Cited by 13SourceScholar
2024

Re-assembling the past: The RePAIR dataset and benchmark for real world 2D and 3D puzzle solving

NeurIPS 2024poster

This paper proposes the RePAIR dataset that represents a challenging benchmark to test modern computational and data driven methods for puzzle-solving and reassembly tasks. Our dataset has unique properties that are uncommon to current benchmarks for 2D and 3D puzzle solving. The fragments and fract…

Cited by 3SourcePDFScholar
2024

SelfGeo: Self-supervised and Geodesic-consistent Estimation of Keypoints on Deformable Shapes

ECCV 2024poster

"Unsupervised 3D keypoints estimation from Point Cloud Data (PCD) is a complex task, even more challenging when an object shape is deforming. As keypoints should be semantically and geometrically consistent across all the 3D frames – each keypoint should be anchored to a specific part of the deformi…

2024

XBG: End-to-End Imitation Learning for Autonomous Behaviour in Human-Robot Interaction and Collaboration

RA-L 2024

This letter presents XBG (eXteroceptive Behaviour Generation), a multimodal end-to-end Imitation Learning (IL) system for whole-body autonomous humanoid robots used in real-world Human-Robot Interaction (HRI) scenarios. The main contribution is an architecture for learning HRI behaviours using a dat

Cited by 6SourcecodeScholar
2023

3DSGrasp: 3D Shape-Completion for Robotic Grasp

ICRA 2023poster

Real-world robotic grasping can be done robustly if a complete 3D Point Cloud Data (PCD) of an object is available. However, in practice, PCDs are often incomplete when objects are viewed from few and sparse viewpoints before the grasping action, leading to the generation of wrong or inaccurate gras…

Cited by 26SourcecodeScholar
2023

Audio-Visual Inpainting: Reconstructing Missing Visual Information with Sound

ICASSP 2023accepted

We tackle audio-visual inpainting, the problem of completing an image in such a way to be consistent with the sound associated to the scene. To this end, we propose a multimodal, audio-visual inpainting method (AVIN), and show how to leverage sound to reconstruct semantically consistent images. AVIN…

Cited by 0SourceScholar
2023

Guiding Pseudo-Labels With Uncertainty Estimation for Source-Free Unsupervised Domain Adaptation

CVPR 2023poster

Standard Unsupervised Domain Adaptation (UDA) methods assume the availability of both source and target data during the adaptation. In this work, we investigate Source-free Unsupervised Domain Adaptation (SF-UDA), a specific case of UDA where a model is adapted to a target domain without access to s…

2023

Person Re-Identification without Identification via Event anonymization

ICCV 2023poster

Wide-scale use of visual surveillance in public spaces puts individual privacy at stake while increasing resource consumption (energy, bandwidth, and computation). Neuromorphic vision sensors (event-cameras) have been recently considered a valid solution to the privacy issue because they do not capt…

Cited by 22PDFcodeScholar
2023

SC3K: Self-supervised and Coherent 3D Keypoints Estimation from Rotated, Noisy, and Decimated Point Cloud Data

ICCV 2023poster

This paper proposes a new method to infer keypoints from arbitrary object categories in practical scenarios where point cloud data (PCD) are noisy, down-sampled and arbitrarily rotated. Our proposed model adheres to the following principles: i) keypoints inference is fully unsupervised (no annotatio…

Cited by 12PDFcodeScholar
2022

Fusion and Orthogonal Projection for Improved Face-Voice Association

ICASSP 2022accepted

We study the problem of learning association between face and voice. Prior works adopt pairwise or triplet loss formulations to learn an embedding space amenable for associated matching and verification tasks. Albeit showing some progress, such loss formulations are restrictive due to dependency on…

Cited by 0SourceScholar
2022

PoserNet: Refining Relative Camera Poses Exploiting Object Detections

ECCV 2022poster

"The estimation of the camera poses associated with a set of images commonly relies on feature matches between the images. In contrast, we are the first to address this challenge by using objectness regions to guide the pose estimation problem rather than explicit semantic object detections. We prop…

2022

Spatial Commonsense Graph for Object Localisation in Partial Scenes

CVPR 2022poster

We solve object localisation in partial scenes, a new problem of estimating the unknown position of an object (e.g. where is the bag?) given a partial 3D scan of a scene. The proposed solution is based on a novel scene graph model, the Spatial Commonsense Graph (SCG), where objects are the nodes and…

Cited by 23PDFcodeScholar
2021

(Just) A Spoonful of Refinements Helps the Registration Error Go Down

ICCV 2021poster

In this paper, we tackle data-driven 3D point cloud registration. Given point correspondences, the standard Kabsch algorithm provides an optimal rotation estimate. This allows to train registration models in an end-to-end manner by differentiating the SVD operation. However, given the initial rotati…

Cited by 3PDFcodeScholar
2021

Audio-Visual Localization by Synthetic Acoustic Image Generation

AAAI 2021technical

Acoustic images constitute an emergent data modality for multimodal scene understanding. Such images have the peculiarity to distinguish the spectral signature of sounds coming from different directions in space, thus providing richer information than the one derived from mono and binaural microphon…

2021

POMP++: Pomcp-based Active Visual Search in unknown indoor environments

IROS 2021poster

In this paper, we focus on the problem of learning online an optimal policy for Active Visual Search (AVS) of objects in unknown indoor environments. We propose POMP++, a planning strategy that introduces a novel formulation on top of the classic Partially Observable Monte Carlo Planning (POMCP) fra…

Cited by 17SourceScholar
2020

Where to Explore Next? ExHistCNN for History-aware Autonomous 3D Exploration

ECCV 2020poster

In this work we address the problem of autonomous 3D exploration of an unknown indoor environment using a depth camera. We cast the problem as the estimation of the Next Best View (NBV) that maximises the coverage of the unknown area. We do this by re-formulating NBV estimation as a classification p…

2019

Autonomous 3-D Reconstruction, Mapping, and Exploration of Indoor Environments With a Robotic Arm

RA-L 2019

We propose a novel information gain metric that combines hand-crafted and data-driven metrics to address the next best view problem for autonomous 3-D mapping of unknown indoor environments. For the hand-crafted metric, we propose an entropy-based information gain that accounts for the previous view

Cited by 48SourceScholar
2018

MX-LSTM: Mixing Tracklets and Vislets to Jointly Forecast Trajectories and Head Poses

CVPR 2018poster

Recent approaches on trajectory forecasting use tracklets to predict the future positions of pedestrians exploiting Long Short Term Memory (LSTM) architectures. This paper shows that adding vislets, that is, short sequences of head pose estimations, allows to increase significantly the trajectory fo…

Cited by 153SourcePDFScholar
2016

Fast 6D pose estimation for texture-less objects from a single RGB image

ICRA 2016

A fundamental step to solve bin-picking and grasping problems is the accurate estimation of an object 3D pose. Such visual task usually rely on profusely textured objects: standard procedures such as detection of interest points or computation of appearance-based descriptors are favoured by using a

Cited by 42SourceScholar
2015

Sparse Representation Classification With Manifold Constraints Transfer

CVPR 2015poster

The fact that image data samples lie on a manifold has been successfully exploited in many learning and inference problems. In this paper we leverage the specific structure of data in order to improve recognition accuracies in general recognition tasks. In particular we propose a novel framework tha…

Cited by 66SourcePDFScholar