← Search

Daichi Azuma

5 accepted papers

2025

CityNav: A Large-Scale Dataset for Real-World Aerial Navigation

ICCV 2025poster

Vision-and-language navigation (VLN) aims to develop agents capable of navigating in realistic environments. While recent cross-modal training approaches have significantly improved navigation performance in both indoor and outdoor scenarios, aerial navigation over real-world cities remains underexp…

Cited by 0SourcePDFScholar
2025

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields

ICCV 2025poster

The advancement of 3D language fields has enabled intuitive interactions with 3D scenes via natural language. However, existing approaches are typically limited to small-scale environments, lacking the scalability and compositional reasoning capabilities necessary for large, complex urban settings.…

Cited by 0SourcePDFScholar
2024

Answerability Fields: Answerable Location Estimation via Diffusion Models

IROS 2024

We propose Answerability Fields (AnsFields), a novel approach for predicting the answerability of questions at different locations within indoor environments. AnsFields is represented as a map, where each grid’s score reflects how well a question can be answered using the panoramic image at that loc

Cited by 0SourceScholar
2024

Map-based Modular Approach for Zero-shot Embodied Question Answering

IROS 2024poster

Embodied Question Answering (EQA) serves as a benchmark task to evaluate the capability of robots to navigate within novel environments and identify objects in response to human queries. However, existing EQA methods often rely on simulated environments and operate with limited vocabularies. This pa…

Cited by 3SourcecodeScholar
2022

ScanQA: 3D Question Answering for Spatial Scene Understanding

CVPR 2022oral

We propose a new 3D spatial understanding task of 3D Question Answering (3D-QA). In the 3D-QA task, models receive visual information from the entire 3D scene of the rich RGB-D indoor scan and answer the given textual questions about the 3D scene. Unlike the 2D-question answering of VQA, the convent…

Cited by 199PDFcodeScholar