← Search

Komei Sugiura

36 accepted papers

2026

Affordance RAG: Hierarchical Multimodal Retrieval With Affordance-Aware Embodied Memory for Mobile Manipulation

RA-L 2026

In this study, we address the problem of openvocabulary mobile manipulation, where a robot is required to carry a wide range of objects to receptacles based on freeform natural language instructions. This task is challenging, as it involves understanding visual semantics and the affordance of manipu

Cited by 2SourceScholar
2026

Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation

ICRA 2026poster

In this study, we address the problem of open-vocabulary mobile manipulation, where a robot is required to carry a wide range of objects to receptacles based on free-form natural language instructions. This task is challenging, as it involves understanding visual semantics and the affordance of mani…

2026

CONDITION-INVARIANT FMRI DECODING OF SPEECH INTELLIGIBILITY WITH DEEP STATE SPACE MODEL

ICASSP 2026poster

Clarifying the neural basis of speech intelligibility is critical for computational neuroscience and digital speech processing. Recent neuroimaging studies have shown that intelligibility modulates cortical activity beyond simple acoustics, primarily in the superior temporal and inferior frontal gyr…

Cited by 0SourcePDFScholar
2026

LILAC: Language-Conditioned Object-Centric Optical Flow for Open-Loop Trajectory Generation

RA-L 2026

We address language-conditioned robotic manipulation using flow-based trajectory generation, which enables training on human and web videos of object manipulation and requires only minimal embodiment-specific data. This task is challenging, as object trajectory generation from pre-manipulation image

Cited by 3SourceScholar
2026

LLM-Free Image Captioning Evaluation in Reference-Flexible Settings

AAAI 2026technical

We focus on the automatic evaluation of image captions in both reference-based and reference-free settings. Existing metrics based on large language models (LLMs) favor their own generations; therefore, the neutrality is in question. Most LLM-free metrics do not suffer from such an issue, whereas th

Cited by 0SourcePDFScholar
2026

Mobile Manipulation Instruction Generation from Multiple Images with Automatic Metric Enhancement

ICRA 2026poster

We consider the problem of generating mobile manipulation instructions based on a target object image and receptacle image. Conventional image captioning models are not able to generate appropriate instructions because their architectures are typically optimized for single-image. In this study, we p…

2026

Open-Vocabulary Mobile Manipulation Based on Double Relaxed Contrastive Learning with Dense Labeling

ICRA 2026poster

Growing labor shortages are increasing the demand for domestic service robots (DSRs) to assist in various settings. In this study, we develop a DSR that transports everyday objects to specified pieces of furniture based on open-vocabulary instructions. Our approach focuses on retrieving images of ta…

2026

ReMoRa: Multimodal Large Language Model based on Refined Motion Representation for Long-Video Understanding

CVPR 2026

While multimodal large language models (MLLMs) have shown remarkable success across a wide range of tasks, long-form video understanding remains a significant challenge.In this study, we focus on video understanding by MLLMs.This task is challenging because processing a full stream of RGB frames is

Cited by 0SourceScholar
2026

ZINA: Multimodal Fine-grained Hallucination Detection and Editing

CVPR 2026

Multimodal Large Language Models (MLLMs) often generate hallucinations, where the output deviates from the visual content. Given that these hallucinations can take diverse forms, detecting hallucinations at a fine-grained level is essential for comprehensive evaluation and analysis. To this end, we

Cited by 0SourcecodeScholar
2025

Deep Space Weather Model: Long-Range Solar Flare Prediction from Multi-Wavelength Images

ICCV 2025poster

Accurate, reliable solar flare prediction is crucial for mitigating potential disruptions to critical infrastructure, while predicting solar flares remains a significant challenge. Existing methods based on heuristic physical features often lack representation learning from solar images. On the othe…

2025

GENNAV: Polygon Mask Generation for Generalized Referring Navigable Regions

CoRL 2025poster

We focus on the task of identifying the location of target regions from a natural language instruction and a front camera image captured by a mobility. This task is challenging because it requires both existence prediction and segmentation mask generation, particularly for stuff-type target regions…

Cited by 0SourceScholar
2025

Interactive Robot Action Replanning using Multimodal LLM Trained from Human Demonstration Videos

ICASSP 2025accepted

Understanding human actions could allow robots to perform a large spectrum of complex manipulation tasks and make collaboration with humans easier. Recently, multimodal scene understanding using audio-visual Transformers has been used to generate robot action sequences from videos of human demonstra…

Cited by 0SourceScholar
2025

Mobile Manipulation Instruction Generation From Multiple Images With Automatic Metric Enhancement

RA-L 2025

We consider the problem of generating free-form mobile manipulation instructions based on a target object image and receptacle image. Conventional image captioning models are not able to generate appropriate instructions because their architectures are typically optimized for single-image. In this s

Cited by 0SourceScholar
2025

Multimodal Target Localization With Landmark-Aware Positioning for Urban Mobility

RA-L 2025

Advancements in vehicle automation technology are expected to significantly impact how humans interact with vehicles. In this study, we propose a method to create user-friendly control interfaces for autonomous vehicles in urban environments. The proposed model predicts the vehicle's destination on

Cited by 1SourceScholar
2025

Open-Vocabulary Mobile Manipulation Based on Double Relaxed Contrastive Learning With Dense Labeling

RA-L 2025

Growing labor shortages are increasing the demand for domestic service robots (DSRs) to assist in various settings. In this study, we develop a DSR that transports everyday objects to specified pieces of furniture based on open-vocabulary instructions. Our approach focuses on retrieving images of ta

Cited by 3SourceScholar
2025

VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions

EMNLP 2025

In this study, we focus on the automatic evaluation of long and detailed image captions generated by multimodal Large Language Models (MLLMs). Most existing automatic evaluation metrics for image captioning are primarily designed for short captions and are not suitable for evaluating long captions.

Cited by 0SourcePDFScholar
2024

Learning-To-Rank Approach for Identifying Everyday Objects Using a Physical-World Search Engine

RA-L 2024

Domestic service robots offer a solution to the increasing demand for daily care and support. A human-in-the-loop approach that combines automation and operator intervention is considered to be a realistic approach to their use in society. Therefore, we focus on the task of retrieving target objects

Cited by 9SourcecodeScholar
2024

Object Segmentation from Open-Vocabulary Manipulation Instructions Based on Optimal Transport Polygon Matching with Multimodal Foundation Models

IROS 2024poster

We consider the task of generating segmentation masks for the target object from an object manipulation instruction, which allows users to give open vocabulary instructions to domestic service robots. Conventional segmentation generation approaches often fail to account for objects outside the camer…

Cited by 1SourceScholar
2024

Polos: Multimodal Metric Learning from Human Feedback for Image Captioning

CVPR 2024highlight

Establishing an automatic evaluation metric that closely aligns with human judgments is essential for effectively developing image captioning models. Recent data-driven metrics have demonstrated a stronger correlation with human judgments than classic metrics such as CIDEr; however they lack suffici…

2024

Task Success Prediction for Open-Vocabulary Manipulation Based on Multi-Level Aligned Representations

CoRL 2024poster

In this study, we consider the problem of predicting task success for open-vocabulary manipulation by a manipulator, based on instruction sentences and egocentric images before and after manipulation. Conventional approaches, including multimodal large language models (MLLMs), often fail to appropri…

Cited by 2SourceScholar
2024

Trimodal Navigable Region Segmentation Model: Grounding Navigation Instructions in Urban Areas

RA-L 2024

In this study, we develop a model that enables mobilities to have more friendly interactions with users. Specifically, we focus on the referring navigable regions task in which a model grounds navigable regions of the road using the mobility's camera image and natural language navigation instruction

Cited by 6SourceScholar
2023

Multimodal Diffusion Segmentation Model for Object Segmentation from Manipulation Instructions

IROS 2023poster

In this study, we aim to develop a model that comprehends a natural language instruction (e.g., “Go to the living room and get the nearest pillow to the radio art on the wall”) and generates a segmentation mask for the target everyday object. The task is challenging because it requires (1) the under…

Cited by 6SourceScholar
2023

Prototypical Contrastive Transfer Learning for Multimodal Language Understanding

IROS 2023poster

Although domestic service robots are expected to assist individuals who require support, they cannot currently interact smoothly with people through natural language. For example, given the instruction “Bring me a bottle from the kitchen,” it is difficult for such robots to specify the bottle in an…

Cited by 2SourceScholar
2023

Switching Head-Tail Funnel UNITER for Dual Referring Expression Comprehension with Fetch-and-Carry Tasks

IROS 2023poster

This paper describes a domestic service robot (DSR) that fetches everyday objects and carries them to specified destinations according to free-form natural language instructions. Given an instruction such as “Move the bottle on the left side of the plate to the empty chair,” the DSR is expected to i…

Cited by 11SourceScholar
2022

Shared Transformer Encoder with Mask-Based 3d Model Estimation for Container Mass Estimation

ICASSP 2022accepted

For human-safe robot control in human-to-robot handover, the physical properties of containers and fillings should be accurately estimated. In this paper, we propose a Transformer encoder that shares the same architecture and parameters for filling level and type estimation. We also propose a mask-b…

Cited by 0SourceScholar
2021

CrossMap Transformer: A Crossmodal Masked Path Transformer Using Double Back-Translation for Vision-and-Language Navigation

RA-L 2021

Navigation guided by natural language instructions is particularly suitable for Domestic Service Robots that interacts naturally with users. This task involves the prediction of a sequence of actions that leads to a specified destination given a natural language navigation instruction. The task thus

Cited by 15SourceScholar
2021

Target-Dependent UNITER: A Transformer-Based Multimodal Language Comprehension Model for Domestic Service Robots

RA-L 2021

Currently, domestic service robots have an insufficient ability to interact naturally through language. This is because understanding human instructions is complicated by various ambiguities. In existing methods, the referring expressions that specify the relationships between objects were insuffici

Cited by 13SourceScholar
2021

Unified Questioner Transformer for Descriptive Question Generation in Goal-Oriented Visual Dialogue

ICCV 2021poster

Building an interactive artificial intelligence that can ask questions about the real world is one of the biggest challenges for vision and language problems. In particular, goal-oriented visual dialogue, where the aim of the agent is to seek information by asking questions during a turn-taking dial…

Cited by 19PDFcodeScholar
2020

A Multimodal Target-Source Classifier With Attention Branches to Understand Ambiguous Instructions for Fetching Daily Objects

RA-L 2020

In this study, we focus on multimodal language understanding for fetching instructions in the domestic service robots context. This task consists of predicting a target object, as instructed by the user, given an image and an unstructured sentence, such as “Bring me the yellow box (from the wooden c

Cited by 10SourceScholar
2020

Alleviating the Burden of Labeling: Sentence Generation by Attention Branch Encoder-Decoder Network

RA-L 2020

Domestic service robots (DSRs) are a promising solution to the shortage of home care workers. However, one of the main limitations of DSRs is their inability to interact naturally through language. Recently, data-driven approaches have been shown to be effective for tackling this limitation; however

Cited by 12SourceScholar
2019

Multimodal Attention Branch Network for Perspective-Free Sentence Generation

CoRL 2019

In this paper, we address the automatic sentence generation of fetching instructions for domestic service robots. Typical fetching commands such as “bring me the yellow toy from the upper part of the white shelf” includes referring expressions, i.e., “from the white upper part of the white shelf”. T

Cited by 0SourcePDFScholar
2019

Understanding Natural Language Instructions for Fetching Daily Objects Using GAN-Based Multimodal Target-Source Classification

RA-L 2019

In this letter, we address multimodal language understanding with unconstrained fetching instruction for domestic service robots. A typical fetching instruction such as “Bring me the yellow toy from the white shelf” requires to infer the user intention, i.e., what object (target) to fetch and from w

Cited by 35SourceScholar
2018

A Multimodal Classifier Generative Adversarial Network for Carry and Place Tasks From Ambiguous Language Instructions

RA-L 2018

This letter focuses on a multimodal language understanding method for carry-and-place tasks with domestic service robots. We address the case of ambiguous instructions, that is, when the target area is not specified. For instance “put away the milk and cereal” is a natural instruction where there is

Cited by 31SourceScholar