← Search

Shuhei Kurita

22 accepted papers

2026

Developing Vision-Language-Action Model from Egocentric Videos

ICRA 2026poster

Egocentric videos capture how humans manipulate objects and tools, providing diverse motion cues for learning object manipulation. Unlike the costly, expert-driven manual teleoperation commonly used in training Vision-Language-Action models (VLAs), egocentric videos offer a scalable alternative. How…

2026

ReMoRa: Multimodal Large Language Model based on Refined Motion Representation for Long-Video Understanding

CVPR 2026

While multimodal large language models (MLLMs) have shown remarkable success across a wide range of tasks, long-form video understanding remains a significant challenge.In this study, we focus on video understanding by MLLMs.This task is challenging because processing a full stream of RGB frames is

Cited by 0SourceScholar
2025

CityNav: A Large-Scale Dataset for Real-World Aerial Navigation

ICCV 2025poster

Vision-and-language navigation (VLN) aims to develop agents capable of navigating in realistic environments. While recent cross-modal training approaches have significantly improved navigation performance in both indoor and outdoor scenarios, aerial navigation over real-world cities remains underexp…

Cited by 0SourcePDFScholar
2025

Constructing Multimodal Datasets from Scratch for Rapid Development of a Japanese Visual Language Model

NAACL 2025system demonstrations

To develop high-performing Visual Language Models (VLMs), it is essential to prepare multimodal resources, such as image-text pairs, interleaved data, and instruction data. While multimodal resources for English are abundant, there is a significant lack of corresponding resources for non-English lan…

2025

Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision

CVPR 2025highlight

Learning to use tools or objects in common scenes, particularly handling them in various ways as instructed, is a key challenge for developing interactive robots. Training models to generate such manipulation trajectories requires a large and diverse collection of detailed manipulation demonstration…

Cited by 0SourcePDFScholar
2025

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields

ICCV 2025poster

The advancement of 3D language fields has enabled intuitive interactions with 3D scenes via natural language. However, existing approaches are typically limited to small-scale environments, lacking the scalability and compositional reasoning capabilities necessary for large, complex urban settings.…

Cited by 0SourcePDFScholar
2025

LegalViz: Legal Text Visualization by Text To Diagram Generation

NAACL 2025long

Legal documents including judgments and court orders require highly sophisticated legal knowledge for understanding. To disclose expert knowledge for non-experts, we explore the problem of visualizing legal texts with easy-to-understand diagrams and propose a novel dataset of LegalViz with 23 langua…

2025

Low-Latency Privacy-Aware Robot Behavior guided by Automatically Generated Text Datasets

IROS 2025

Humans typically avert their gaze when faced with situations involving another person’s privacy, and humanoid robots should exhibit similar behaviors. Various approaches exist for privacy recognition, including an image privacy recognition model and a Large Vision-Language Model (LVLM). The former r

Cited by 0SourceScholar
2025

Referring Expression Comprehension for Small Objects

ICCV 2025poster

Referring expression comprehension (REC) aims to localize the target object described by a natural language expression.Recent advances in vision-language learning have led to significant performance improvements in REC tasks.However, localizing extremely small objects remains a considerable challeng…

2024

Answerability Fields: Answerable Location Estimation via Diffusion Models

IROS 2024

We propose Answerability Fields (AnsFields), a novel approach for predicting the answerability of questions at different locations within indoor environments. AnsFields is represented as a map, where each grid’s score reflects how well a question can be answered using the panoramic image at that loc

Cited by 0SourceScholar
2024

JDocQA: Japanese Document Question Answering Dataset for Generative Language Models

COLING 2024main

Document question answering is a task of question answering on given documents such as reports, slides, pamphlets, and websites, and it is a truly demanding task as paper and electronic forms of documents are so common in our society. This is known as a quite challenging task because it requires not…

2024

Map-based Modular Approach for Zero-shot Embodied Question Answering

IROS 2024poster

Embodied Question Answering (EQA) serves as a benchmark task to evaluate the capability of robots to navigate within novel environments and identify objects in response to human queries. However, existing EQA methods often rely on simulated environments and operate with limited vocabularies. This pa…

Cited by 3SourcecodeScholar
2024

Text360Nav: 360-Degree Image Captioning Dataset for Urban Pedestrians Navigation

COLING 2024main

Text feedback from urban scenes is a crucial tool for pedestrians to understand surroundings, obstacles, and safe pathways. However, existing image captioning datasets often concentrate on the overall image description and lack detailed scene descriptions, overlooking features for pedestrians walkin…

2023

ARKitSceneRefer: Text-based Localization of Small Objects in Diverse Real-World 3D Indoor Scenes

EMNLP 2023long findings

3D referring expression comprehension is a task to ground text representations onto objects in 3D scenes. It is a crucial task for indoor household robots or augmented reality devices to localize objects referred to in user instructions. However, existing indoor 3D referring expression comprehension…

Cited by 0SourceScholar
2023

CityRefer: Geography-aware 3D Visual Grounding Dataset on City-scale Point Cloud Data

NeurIPS 2023poster

City-scale 3D point cloud is a promising way to express detailed and complicated outdoor structures. It encompasses both the appearance and geometry features of segmented city components, including cars, streets, and buildings that can be utilized for attractive applications such as user-interactive…

2023

Query-based Image Captioning from Multi-context 360° Images

EMNLP 2023long findings

A 360-degree image captures the entire scene without the limitations of a camera's field of view, which makes it difficult to describe all the contexts in a single caption. We propose a novel task called Query-based Image Captioning (QuIC) for 360-degree images, where a query (words or short phrases…

Cited by 1SourceScholar
2023

RefEgo: Referring Expression Comprehension Dataset from First-Person Perception of Ego4D

ICCV 2023poster

Grounding textual expressions on scene objects from first-person views is a truly demanding capability in developing agents that are aware of their surroundings and behave following intuitive text instructions. Such capability is of necessity for glass-devices or autonomous robots to localize referr…

Cited by 17PDFcodeScholar
2022

Iterative Span Selection: Self-Emergence of Resolving Orders in Semantic Role Labeling

COLING 2022main

Semantic Role Labeling (SRL) is the task of labeling semantic arguments for marked semantic predicates. Semantic arguments and their predicates are related in various distinct manners, of which certain semantic arguments are a necessity while others serve as an auxiliary to their predicates. To cons…

Cited by 2SourcePDFScholar
2022

ScanQA: 3D Question Answering for Spatial Scene Understanding

CVPR 2022oral

We propose a new 3D spatial understanding task of 3D Question Answering (3D-QA). In the 3D-QA task, models receive visual information from the entire 3D scene of the rich RGB-D indoor scan and answer the given textual questions about the 3D scene. Unlike the 2D-question answering of VQA, the convent…

Cited by 199PDFcodeScholar
2022

Visual Recipe Flow: A Dataset for Learning Visual State Changes of Objects with Recipe Flows

COLING 2022main

We present a new multimodal dataset called Visual Recipe Flow, which enables us to learn a cooking action result for each object in a recipe text. The dataset consists of object state changes and the workflow of the recipe text. The state change is represented as an image pair, while the workflow is…

Cited by 11SourcePDFScholar
2021

Co-Teaching Student-Model through Submission Results of Shared Task

EMNLP 2021finding

Shared tasks have a long history and have become the mainstream of NLP research. Most of the shared tasks require participants to submit only system outputs and descriptions. It is uncommon for the shared task to request submission of the system itself because of the license issues and implementatio…

2021

Generative Language-Grounded Policy in Vision-and-Language Navigation with Bayes' Rule

ICLR 2021poster

Vision-and-language navigation (VLN) is a task in which an agent is embodied in a realistic 3D environment and follows an instruction to reach the goal node. While most of the previous studies have built and investigated a discriminative approach, we notice that there are in fact two possible approa…

Cited by 27SourcePDFScholar