← Search

Tomoya Yoshida

2 accepted papers

2026

Developing Vision-Language-Action Model from Egocentric Videos

ICRA 2026poster

Egocentric videos capture how humans manipulate objects and tools, providing diverse motion cues for learning object manipulation. Unlike the costly, expert-driven manual teleoperation commonly used in training Vision-Language-Action models (VLAs), egocentric videos offer a scalable alternative. How…

2025

Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision

CVPR 2025highlight

Learning to use tools or objects in common scenes, particularly handling them in various ways as instructed, is a key challenge for developing interactive robots. Training models to generate such manipulation trajectories requires a large and diverse collection of detailed manipulation demonstration…

Cited by 0SourcePDFScholar