← Search

Taiki Miyanishi

10 accepted papers

2025

CityNav: A Large-Scale Dataset for Real-World Aerial Navigation

ICCV 2025poster

Vision-and-language navigation (VLN) aims to develop agents capable of navigating in realistic environments. While recent cross-modal training approaches have significantly improved navigation performance in both indoor and outdoor scenarios, aerial navigation over real-world cities remains underexp…

Cited by 0SourcePDFScholar
2025

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields

ICCV 2025poster

The advancement of 3D language fields has enabled intuitive interactions with 3D scenes via natural language. However, existing approaches are typically limited to small-scale environments, lacking the scalability and compositional reasoning capabilities necessary for large, complex urban settings.…

Cited by 0SourcePDFScholar
2025

LegalViz: Legal Text Visualization by Text To Diagram Generation

NAACL 2025long

Legal documents including judgments and court orders require highly sophisticated legal knowledge for understanding. To disclose expert knowledge for non-experts, we explore the problem of visualizing legal texts with easy-to-understand diagrams and propose a novel dataset of LegalViz with 23 langua…

2024

Answerability Fields: Answerable Location Estimation via Diffusion Models

IROS 2024

We propose Answerability Fields (AnsFields), a novel approach for predicting the answerability of questions at different locations within indoor environments. AnsFields is represented as a map, where each grid’s score reflects how well a question can be answered using the panoramic image at that loc

Cited by 0SourceScholar
2024

JDocQA: Japanese Document Question Answering Dataset for Generative Language Models

COLING 2024main

Document question answering is a task of question answering on given documents such as reports, slides, pamphlets, and websites, and it is a truly demanding task as paper and electronic forms of documents are so common in our society. This is known as a quite challenging task because it requires not…

2024

Map-based Modular Approach for Zero-shot Embodied Question Answering

IROS 2024poster

Embodied Question Answering (EQA) serves as a benchmark task to evaluate the capability of robots to navigate within novel environments and identify objects in response to human queries. However, existing EQA methods often rely on simulated environments and operate with limited vocabularies. This pa…

Cited by 3SourcecodeScholar
2023

CityRefer: Geography-aware 3D Visual Grounding Dataset on City-scale Point Cloud Data

NeurIPS 2023poster

City-scale 3D point cloud is a promising way to express detailed and complicated outdoor structures. It encompasses both the appearance and geometry features of segmented city components, including cars, streets, and buildings that can be utilized for attractive applications such as user-interactive…

2023

Query-based Image Captioning from Multi-context 360° Images

EMNLP 2023long findings

A 360-degree image captures the entire scene without the limitations of a camera's field of view, which makes it difficult to describe all the contexts in a single caption. We propose a novel task called Query-based Image Captioning (QuIC) for 360-degree images, where a query (words or short phrases…

Cited by 1SourceScholar
2022

ScanQA: 3D Question Answering for Spatial Scene Understanding

CVPR 2022oral

We propose a new 3D spatial understanding task of 3D Question Answering (3D-QA). In the 3D-QA task, models receive visual information from the entire 3D scene of the rich RGB-D indoor scan and answer the given textual questions about the 3D scene. Unlike the 2D-question answering of VQA, the convent…

Cited by 199PDFcodeScholar
2018

A Sparse Coding Framework for Gaze Prediction in Egocentric Video

ICASSP 2018accepted

To efficiently process and understand a large amount of incoming visual information from first-person perspective (i.e. egocentric vision), predicting human gaze is important. However, even though people continuously gaze in noisy environments, most existing gaze prediction methods mainly use image…

Cited by 0SourceScholar