← Search

Shinsuke Mori

5 accepted papers

2026

Developing Vision-Language-Action Model from Egocentric Videos

ICRA 2026poster

Egocentric videos capture how humans manipulate objects and tools, providing diverse motion cues for learning object manipulation. Unlike the costly, expert-driven manual teleoperation commonly used in training Vision-Language-Action models (VLAs), egocentric videos offer a scalable alternative. How…

2025

Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision

CVPR 2025highlight

Learning to use tools or objects in common scenes, particularly handling them in various ways as instructed, is a key challenge for developing interactive robots. Training models to generate such manipulation trajectories requires a large and diverse collection of detailed manipulation demonstration…

Cited by 0SourcePDFScholar
2024

Automatic Construction of a Large-Scale Corpus for Geoparsing Using Wikipedia Hyperlinks

COLING 2024main

Geoparsing is the task of estimating the latitude and longitude (coordinates) of location expressions in texts. Geoparsing must deal with the ambiguity of the expressions that indicate multiple locations with the same notation. For evaluating geoparsing systems, several corpora have been proposed in…

Cited by 0SourcePDFScholar
2024

Vision-Language Interpreter for Robot Task Planning

ICRA 2024poster

Large language models (LLMs) are accelerating the development of language-guided robot planners. Meanwhile, symbolic planners offer the advantage of interpretability. This paper proposes a new task that bridges these two trends, namely, multimodal planning problem specification. The aim is to genera…

Cited by 62SourcecodeScholar
2022

Visual Recipe Flow: A Dataset for Learning Visual State Changes of Objects with Recipe Flows

COLING 2022main

We present a new multimodal dataset called Visual Recipe Flow, which enables us to learn a cooking action result for each object in a recipe text. The dataset consists of object state changes and the workflow of the recipe text. The state change is represented as an image pair, while the workflow is…

Cited by 11SourcePDFScholar