← Search

Yui Iioka

5 accepted papers

2025

GENNAV: Polygon Mask Generation for Generalized Referring Navigable Regions

CoRL 2025poster

We focus on the task of identifying the location of target regions from a natural language instruction and a front camera image captured by a mobility. This task is challenging because it requires both existence prediction and segmentation mask generation, particularly for stuff-type target regions…

Cited by 0SourceScholar
2025

Multimodal Target Localization With Landmark-Aware Positioning for Urban Mobility

RA-L 2025

Advancements in vehicle automation technology are expected to significantly impact how humans interact with vehicles. In this study, we propose a method to create user-friendly control interfaces for autonomous vehicles in urban environments. The proposed model predicts the vehicle's destination on

Cited by 1SourceScholar
2024

Trimodal Navigable Region Segmentation Model: Grounding Navigation Instructions in Urban Areas

RA-L 2024

In this study, we develop a model that enables mobilities to have more friendly interactions with users. Specifically, we focus on the referring navigable regions task in which a model grounds navigable regions of the road using the mobility's camera image and natural language navigation instruction

Cited by 6SourceScholar
2023

Multimodal Diffusion Segmentation Model for Object Segmentation from Manipulation Instructions

IROS 2023poster

In this study, we aim to develop a model that comprehends a natural language instruction (e.g., “Go to the living room and get the nearest pillow to the radio art on the wall”) and generates a segmentation mask for the target everyday object. The task is challenging because it requires (1) the under…

Cited by 6SourceScholar
2022

Shared Transformer Encoder with Mask-Based 3d Model Estimation for Container Mass Estimation

ICASSP 2022accepted

For human-safe robot control in human-to-robot handover, the physical properties of containers and fillings should be accurately estimated. In this paper, we propose a Transformer encoder that shares the same architecture and parameters for filling level and type estimation. We also propose a mask-b…

Cited by 0SourceScholar