← Search

Motonari Kambara

9 accepted papers

2026

LILAC: Language-Conditioned Object-Centric Optical Flow for Open-Loop Trajectory Generation

RA-L 2026

We address language-conditioned robotic manipulation using flow-based trajectory generation, which enables training on human and web videos of object manipulation and requires only minimal embodiment-specific data. This task is challenging, as object trajectory generation from pre-manipulation image

Cited by 3SourceScholar
2026

Mobile Manipulation Instruction Generation from Multiple Images with Automatic Metric Enhancement

ICRA 2026poster

We consider the problem of generating mobile manipulation instructions based on a target object image and receptacle image. Conventional image captioning models are not able to generate appropriate instructions because their architectures are typically optimized for single-image. In this study, we p…

2025

Interactive Robot Action Replanning using Multimodal LLM Trained from Human Demonstration Videos

ICASSP 2025accepted

Understanding human actions could allow robots to perform a large spectrum of complex manipulation tasks and make collaboration with humans easier. Recently, multimodal scene understanding using audio-visual Transformers has been used to generate robot action sequences from videos of human demonstra…

Cited by 0SourceScholar
2025

Mobile Manipulation Instruction Generation From Multiple Images With Automatic Metric Enhancement

RA-L 2025

We consider the problem of generating free-form mobile manipulation instructions based on a target object image and receptacle image. Conventional image captioning models are not able to generate appropriate instructions because their architectures are typically optimized for single-image. In this s

Cited by 0SourceScholar
2024

Learning-To-Rank Approach for Identifying Everyday Objects Using a Physical-World Search Engine

RA-L 2024

Domestic service robots offer a solution to the increasing demand for daily care and support. A human-in-the-loop approach that combines automation and operator intervention is considered to be a realistic approach to their use in society. Therefore, we focus on the task of retrieving target objects

Cited by 9SourcecodeScholar
2024

Object Segmentation from Open-Vocabulary Manipulation Instructions Based on Optimal Transport Polygon Matching with Multimodal Foundation Models

IROS 2024poster

We consider the task of generating segmentation masks for the target object from an object manipulation instruction, which allows users to give open vocabulary instructions to domestic service robots. Conventional segmentation generation approaches often fail to account for objects outside the camer…

Cited by 1SourceScholar
2024

Task Success Prediction for Open-Vocabulary Manipulation Based on Multi-Level Aligned Representations

CoRL 2024poster

In this study, we consider the problem of predicting task success for open-vocabulary manipulation by a manipulator, based on instruction sentences and egocentric images before and after manipulation. Conventional approaches, including multimodal large language models (MLLMs), often fail to appropri…

Cited by 2SourceScholar
2023

Switching Head-Tail Funnel UNITER for Dual Referring Expression Comprehension with Fetch-and-Carry Tasks

IROS 2023poster

This paper describes a domestic service robot (DSR) that fetches everyday objects and carries them to specified destinations according to free-form natural language instructions. Given an instruction such as “Move the bottle on the left side of the plate to the empty chair,” the DSR is expected to i…

Cited by 11SourceScholar