← Search

Kei Katsumata

3 accepted papers

2026

Mobile Manipulation Instruction Generation from Multiple Images with Automatic Metric Enhancement

ICRA 2026poster

We consider the problem of generating mobile manipulation instructions based on a target object image and receptacle image. Conventional image captioning models are not able to generate appropriate instructions because their architectures are typically optimized for single-image. In this study, we p…

2025

GENNAV: Polygon Mask Generation for Generalized Referring Navigable Regions

CoRL 2025poster

We focus on the task of identifying the location of target regions from a natural language instruction and a front camera image captured by a mobility. This task is challenging because it requires both existence prediction and segmentation mask generation, particularly for stuff-type target regions…

Cited by 0SourceScholar
2025

Mobile Manipulation Instruction Generation From Multiple Images With Automatic Metric Enhancement

RA-L 2025

We consider the problem of generating free-form mobile manipulation instructions based on a target object image and receptacle image. Conventional image captioning models are not able to generate appropriate instructions because their architectures are typically optimized for single-image. In this s

Cited by 0SourceScholar