← Search

Angel X Chang

35 accepted papers

2026

Artiverse: A Diverse and Physically Grounded Dataset for Articulated Objects

CVPR 2026

We present Artiverse, a diverse and physically grounded dataset of high-quality articulated 3D objects designed for realistic functional modeling and simulation. Artiverse contains 5.4K human-authored objects across a broad range of 88 categories, aggregated from multiple 3D static repositories. Obj

Cited by 0SourceScholar
2026

Do 3D Large Language Models Really Understand 3D Spatial Relationships?

ICLR 2026poster

Recent 3D Large-Language Models (3D-LLMs) claim to understand 3D worlds, especially spatial relationships among objects. Yet, we find that simply fine-tuning a language model on text-only question-answer pairs can perform comparably or even surpass these methods on the SQA3D benchmark without using…

Cited by 0SourceScholar
2026

JRM: Joint Reconstruction Model for Multiple Objects without Alignment

CVPR 2026

Object-centric reconstruction seeks to recover the 3D structure of a scene through composition of independent objects. While this independence can simplify modeling, it discards strong signals that could improve reconstruction, notably repetition where the same object model is seen multiple times in

Cited by 0SourceScholar
2026

ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning

ICML 2026poster

Current evaluations of spatial intelligence can be systematically invalid under modern vision-language model (VLM) settings. First, many benchmarks derive question-answer (QA) pairs from point-cloud-based 3D annotations originally curated for traditional 3D perception. When such annotations are trea…

Cited by 0SourceScholar
2025

CLIBD: Bridging Vision and Genomics for Biodiversity Monitoring at Scale

ICLR 2025poster

Measuring biodiversity is crucial for understanding ecosystem health. While prior works have developed machine learning models for taxonomic classification of photographic images and DNA separately, in this work, we introduce a multi-modal approach combining both, using CLIP-style contrastive learni…

2025

Diorama: Unleashing Zero-shot Single-view 3D Indoor Scene Modeling

ICCV 2025poster

Reconstructing structured 3D scenes from RGB images using CAD objects unlocks efficient and compact scene representations that maintain compositionality and interactability. Existing works propose training-heavy methods relying on either expensive yet inaccurate real-world annotations or controllabl…

Cited by 0SourcePDFScholar
2025

NuiScene: Exploring Efficient Generation of Unbounded Outdoor Scenes

ICCV 2025poster

In this paper, we explore the task of generating expansive outdoor scenes, ranging from castles to high-rises. Unlike indoor scene generation, which has been a primary focus of prior work, outdoor scene generation presents unique challenges, including wide variations in scene heights and the need fo…

2025

SINGAPO: Single Image Controlled Generation of Articulated Parts in Objects

ICLR 2025poster

We address the challenge of creating 3D assets for household articulated objects from a single image. Prior work on articulated object creation either requires multi-view multi-state input, or only allows coarse control over the generation process. These limitations hinder the scalability and practi…

Cited by 6SourcePDFScholar
2024

BIOSCAN-5M: A Multimodal Dataset for Insect Biodiversity

NeurIPS 2024poster

As part of an ongoing worldwide effort to comprehend and monitor insect biodiversity, this paper presents the BIOSCAN-5M Insect dataset to the machine learning community and establish several benchmark tasks. BIOSCAN-5M is a comprehensive dataset containing multi-modal information for over 5 million…

2024

Habitat Synthetic Scenes Dataset (HSSD-200): An Analysis of 3D Scene Scale and Realism Tradeoffs for ObjectGoal Navigation

CVPR 2024poster

We contribute the Habitat Synthetic Scene Dataset a dataset of 211 high-quality 3D scenes and use it to test navigation agent generalization to realistic 3D environments. Our dataset represents real interiors and contains a diverse set of 18656 models of real-world objects. We investigate the impact…

Cited by 53SourcePDFScholar
2024

R3DS: Reality-linked 3D Scenes for Panoramic Scene Understanding

ECCV 2024poster

"We introduce the () dataset of synthetic 3D scenes mirroring the real-world scene arrangements from Matterport3D panoramas. Compared to prior work, has more complete and densely populated scenes with objects linked to real-world observations in panoramas. also provides an object support hierarchy,…

Cited by 1SourcePDFScholar
2023

A Step Towards Worldwide Biodiversity Assessment: The BIOSCAN-1M Insect Dataset

NeurIPS 2023poster

In an effort to catalog insect biodiversity, we propose a new large dataset of hand-labelled insect images, the BIOSCAN-1M Insect Dataset. Each record is taxonomically classified by an expert, and also has associated genetic information including raw nucleotide barcode sequences and assigned barcode…

2023

Exploiting Proximity-Aware Tasks for Embodied Social Navigation

ICCV 2023poster

Learning how to navigate among humans in an occluded and spatially constrained indoor environment, is a key ability required to embodied agents to be integrated into our society. In this paper, we propose an end-to-end architecture that exploits Proximity-Aware Tasks (referred as to Risk and Proximi…

Cited by 15PDFcodeScholar
2023

HomeRobot: Open-Vocabulary Mobile Manipulation

CoRL 2023poster

HomeRobot (noun): An affordable compliant robot that navigates homes and manipulates a wide range of objects in order to complete everyday tasks. Open-Vocabulary Mobile Manipulation (OVMM) is the problem of picking any object in any unseen environment, and placing it in a commanded location. This i…

Cited by 98SourcecodeScholar
2023

UniT3D: A Unified Transformer for 3D Dense Captioning and Visual Grounding

ICCV 2023poster

Performing 3D dense captioning and visual grounding requires a common and shared understanding of the underlying multimodal relationships. However, despite some previous attempts on connecting these two related tasks with highly task-specific neural modules, it remains understudied how to explicitly…

Cited by 89PDFScholar
2022

D3Net: A Unified Speaker-Listener Architecture for 3D Dense Captioning and Visual Grounding

ECCV 2022poster

"Recent work on dense captioning and visual grounding in 3D have achieved impressive results. Despite developments in both areas, the limited amount of available 3D vision-language data causes overfitting issues for 3D visual grounding and 3D dense captioning methods. Also, how to discriminatively d…

Cited by 38SourcePDFScholar
2022

MultiScan: Scalable RGBD scanning for 3D environments with articulated objects

NeurIPS 2022accept

We introduce MultiScan, a scalable RGBD dataset construction pipeline leveraging commodity mobile devices to scan indoor scenes with articulated objects and web-based semantic annotation interfaces to efficiently annotate object and part semantics and part mobility parameters. We use this pipeline t…

Cited by 30SourcePDFScholar
2021

Habitat 2.0: Training Home Assistants to Rearrange their Habitat

NeurIPS 2021spotlight

We introduce Habitat 2.0 (H2.0), a simulation platform for training virtual robots in interactive 3D environments and complex physics-enabled scenarios. We make comprehensive contributions to all levels of the embodied AI stack – data, simulation, and benchmark tasks. Specifically, we present: (i) R…

2021

Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI

NeurIPS 2021poster

We present the Habitat-Matterport 3D (HM3D) dataset. HM3D is a large-scale dataset of 1,000 building-scale 3D reconstructions from a diverse set of real-world locations. Each scene in the dataset consists of a textured 3D mesh reconstruction of interiors such as multi-floor residences, stores, and ot…

Cited by 434SourcecodeScholar
2021

Interpretation of Emergent Communication in Heterogeneous Collaborative Embodied Agents

ICCV 2021poster

Communication between embodied AI agents has received increasing attention in recent years. Despite its use, it is still unclear whether the learned communication is interpretable and grounded in perception. To study the grounding of emergent forms of communication, we first introduce the collaborat…

Cited by 39PDFScholar
2021

Plan2Scene: Converting Floorplans to 3D Scenes

CVPR 2021poster

We address the task of converting a floorplan and a set of associated photos of a residence into a textured 3D mesh model, a task which we call Plan2Scene. Our system 1) lifts a floorplan image to a 3D mesh model; 2) synthesizes surface textures based on the input photos; and 3) infers textures for…

Cited by 32PDFcodeScholar
2020

SAPIEN: A SimulAted Part-Based Interactive ENvironment

CVPR 2020oral

Building home assistant robots has long been a goal for vision and robotics researchers. To achieve this task, a simulated environment with physically realistic simulation, sufficient articulated objects, and transferability to the real robot is indispensable. Existing environments achieve these req…

Cited by 560PDFcodeScholar
2020

ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language

ECCV 2020poster

We introduce the new task of 3D object localization in RGB-D scans using natural language descriptions. As input, we assume a point cloud of a scanned 3D scene along with a free-form description of a specified target object. To address this task, we propose ScanRefer, where the core idea is to learn…

Cited by 413SourcePDFScholar
2019

Hierarchy Denoising Recursive Autoencoders for 3D Scene Layout Prediction

CVPR 2019poster

Indoor scenes exhibit rich hierarchical structure in 3D object layouts. Many tasks in 3D scene understanding can benefit from reasoning jointly about the hierarchical context of a scene, and the identities of objects. We present a variational denoising recursive autoencoder (VDRAE) that generates an…

Cited by 30PDFScholar
2019

PartNet: A Large-Scale Benchmark for Fine-Grained and Hierarchical Part-Level 3D Object Understanding

CVPR 2019poster

We present PartNet: a consistent, large-scale dataset of 3D objects annotated with fine-grained, instance-level, and hierarchical 3D part information. Our dataset consists of 573,585 part instances over 26,671 3D models covering 24 object categories. This dataset enables and serves as a catalyst for…

Cited by 849PDFScholar
2019

Scan2CAD: Learning CAD Model Alignment in RGB-D Scans

CVPR 2019oral

We present Scan2CAD, a novel data-driven method that learns to align clean 3D CAD models from a shape database to the noisy and incomplete geometry of a commodity RGB-D scan. For a 3D reconstruction of an indoor scene, our method takes as input a set of CAD models, and predicts a 9DoF pose that alig…

Cited by 294PDFScholar
2018

Im2Pano3D: Extrapolating 360° Structure and Semantics Beyond the Field of View

CVPR 2018poster

We present Im2Pano3D, a convolutional neural network that generates a dense prediction of 3D structure and a probability distribution of semantic labels for a full 360 panoramic view of an indoor scene when given only a partial observation ( <=50%) in the form of an RGB-D image. To make this possibl…

2017

ScanNet: Richly-Annotated 3D Reconstructions of Indoor Scenes

CVPR 2017spotlight

A key requirement for leveraging supervised deep learning methods is the availability of large, labeled datasets. Unfortunately, in the context of RGB-D scene understanding, very little data is available -- current datasets cover a small range of scene views and have limited semantic annotations.…

Cited by 5003PDFScholar
2017

Semantic Scene Completion From a Single Depth Image

CVPR 2017oral

This paper focuses on semantic scene completion, a task for producing a complete 3D voxel representation of volumetric occupancy and semantic labels for a scene from a single-view depth map observation. Previous work has considered scene completion and semantic labeling of depth maps separately. How…

Cited by 1504PDFcodeScholar