← Search

Daichi Yashima

5 accepted papers

2026

Mobile Manipulation Instruction Generation from Multiple Images with Automatic Metric Enhancement

ICRA 2026poster

We consider the problem of generating mobile manipulation instructions based on a target object image and receptacle image. Conventional image captioning models are not able to generate appropriate instructions because their architectures are typically optimized for single-image. In this study, we p…

2026

Open-Vocabulary Mobile Manipulation Based on Double Relaxed Contrastive Learning with Dense Labeling

ICRA 2026poster

Growing labor shortages are increasing the demand for domestic service robots (DSRs) to assist in various settings. In this study, we develop a DSR that transports everyday objects to specified pieces of furniture based on open-vocabulary instructions. Our approach focuses on retrieving images of ta…

2026

ReMoRa: Multimodal Large Language Model based on Refined Motion Representation for Long-Video Understanding

CVPR 2026

While multimodal large language models (MLLMs) have shown remarkable success across a wide range of tasks, long-form video understanding remains a significant challenge.In this study, we focus on video understanding by MLLMs.This task is challenging because processing a full stream of RGB frames is

Cited by 0SourceScholar
2025

Mobile Manipulation Instruction Generation From Multiple Images With Automatic Metric Enhancement

RA-L 2025

We consider the problem of generating free-form mobile manipulation instructions based on a target object image and receptacle image. Conventional image captioning models are not able to generate appropriate instructions because their architectures are typically optimized for single-image. In this s

Cited by 0SourceScholar
2025

Open-Vocabulary Mobile Manipulation Based on Double Relaxed Contrastive Learning With Dense Labeling

RA-L 2025

Growing labor shortages are increasing the demand for domestic service robots (DSRs) to assist in various settings. In this study, we develop a DSR that transports everyday objects to specified pieces of furniture based on open-vocabulary instructions. Our approach focuses on retrieving images of ta

Cited by 3SourceScholar