← Search

Iro Armeni

20 accepted papers

2026

Do 3D Large Language Models Really Understand 3D Spatial Relationships?

ICLR 2026poster

Recent 3D Large-Language Models (3D-LLMs) claim to understand 3D worlds, especially spatial relationships among objects. Yet, we find that simply fine-tuning a language model on text-only question-answer pairs can perform comparably or even surpass these methods on the SQA3D benchmark without using…

Cited by 0SourceScholar
2026

GaussFusion: Improving 3D Reconstruction in the Wild with A Geometry-Informed Video Generator

CVPR 2026

We present GaussFusion, a novel approach for improving 3D Gaussian splatting (3DGS) reconstructions in the wild through geometry-informed video generation. GaussFusion mitigates common 3DGS artifacts, including floaters, flickering, and blur caused by camera pose errors, incomplete coverage, and noi

Cited by 0SourceScholar
2026

ReScene4D: Temporally Consistent Semantic Instance Segmentation of Evolving Indoor 3D Scenes

CVPR 2026

Indoor environments evolve as objects move, appear, or leave the scene. Capturing these dynamics requires maintaining temporally consistent instance identities across intermittently captured 3D scans, even when changes are unobserved. We introduce and formalize the task of temporally sparse 4D indoo

Cited by 0SourceScholar
2025

CrossOver: 3D Scene Cross-Modal Alignment

CVPR 2025highlight

Multi-modal 3D object understanding has gained significant attention, yet current approaches often assume complete data availability and rigid alignment across all modalities. We present CrossOver, a novel framework for cross-modal 3D scene understanding via flexible, scene-level modality alignment.…

2025

GuideFlow3D: Optimization-Guided Rectified Flow For Appearance Transfer

NeurIPS 2025poster

Transferring appearance to 3D assets using different representations of the appearance object - such as images or text - has garnered interest due to its wide range of applications in industries like gaming, augmented reality, and digital content creation. However, state-of-the-art methods still fai…

Cited by 0SourcecodeScholar
2025

Rectified Point Flow: Generic Point Cloud Pose Estimation

NeurIPS 2025spotlight

We present Rectified Point Flow, a unified parameterization that formulates pairwise point cloud registration and multi-part shape assembly as a single conditional generative problem. Given unposed point clouds, our method learns a continuous point-wise velocity field that transports noisy points to…

Cited by 0SourcecodeScholar
2025

WildGS-SLAM: Monocular Gaussian Splatting SLAM in Dynamic Environments

CVPR 2025poster

We present WildGS-SLAM, a robust and efficient monocular RGB SLAM system designed to handle dynamic environments by leveraging uncertainty-aware geometric mapping. Unlike traditional SLAM systems, which assume static scenes, our approach integrates depth and uncertainty information to enhance tracki…

2024

Living Scenes: Multi-object Relocalization and Reconstruction in Changing 3D Environments

CVPR 2024highlight

Research into dynamic 3D scene understanding has primarily focused on short-term change tracking from dense observations while little attention has been paid to long-term changes with sparse observations. We address this gap with MoRE a novel approach for multi-object relocalization and reconstructi…

2024

Multiway Point Cloud Mosaicking with Diffusion and Global Optimization

CVPR 2024poster

We introduce a novel framework for multiway point cloud mosaicking (named Wednesday) designed to co-align sets of partially overlapping point clouds -- typically obtained from 3D scanners or moving RGB-D cameras -- into a unified coordinate system. At the core of our approach is ODIN a learned pairw…

2023

Learning-based Relational Object Matching Across Views

ICRA 2023poster

Intelligent robots require object-level scene understanding to reason about possible tasks and interactions with the environment. Moreover, many perception tasks such as scene reconstruction, image retrieval, or place recognition can benefit from reasoning on the level of objects. While keypoint-bas…

Cited by 5SourceScholar
2023

SGAligner: 3D Scene Alignment with Scene Graphs

ICCV 2023poster

Building 3D scene graphs has recently emerged as a topic in scene representation for several embodied AI applications to represent the world in a structured and rich manner. With their increased use in solving downstream tasks (e.g., navigation and room rearrangement), can we leverage and recycle th…

Cited by 14PDFcodeScholar
2020

Robust Policies via Mid-Level Visual Representations: An Experimental Study in Manipulation and Navigation

CoRL 2020

Vision-based robotics often factors the control loop into separate components for perception and control. Conventional perception components usually extract hand-engineered features from the visual input that are then used by the control component in an explicit manner. In contrast, recent advances

Cited by 0SourcePDFScholar
2019

3D Scene Graph: A Structure for Unified Semantics, 3D Space, and Camera

ICCV 2019poster

A comprehensive semantic understanding of a scene is important for many applications - but in what space should diverse semantic information (e.g., objects, scene categories, material types, 3D shapes, etc.) be grounded and what should be its structure? Aspiring to have one unified structure that ho…

Cited by 413PDFcodeScholar
2016

3D Semantic Parsing of Large-Scale Indoor Spaces

CVPR 2016oral

In this paper, we propose a method for semantic parsing the 3D point cloud of an entire building using a hierarchical approach: first, the raw data is parsed into semantically meaningful spaces (e.g. rooms, etc) that are aligned into a canonical reference coordinate system. Second, the spaces are pa…

Cited by 2242PDFScholar