← Search

Naoki Wake

4 accepted papers

2024

GPT-4V(ision) for Robotics: Multimodal Task Planning From Human Demonstration

RA-L 2024

We introduce a pipeline that enhances a general-purpose Vision Language Model, GPT-4V(ision), to facilitate one-shot visual teaching for robotic manipulation. This system analyzes videos of humans performing tasks and outputs executable robot programs that incorporate insights into affordances. The

Cited by 117SourceScholar
2022

Deep Gesture Generation for Social Robots Using Type-Specific Libraries

IROS 2022poster

Body language such as conversational gesture is a powerful way to ease communication. Conversational gestures do not only make a speech more lively but also contain semantic meaning that helps to stress important information in the discussion. In the field of robotics, giving conversational agents (…

Cited by 6SourceScholar
2021

Task-Oriented Motion Mapping on Robots of Various Configuration Using Body Role Division

RA-L 2021

Many works in robot teaching either focus only on teaching task knowledge, such as geometric constraints, or motion knowledge, such as the motion for accomplishing a task. However, to effectively teach a complex task sequence to a robot, it is important to take advantage of both task and motion know

Cited by 24SourceScholar
2021

Verbal Focus-of-Attention System for Learning-from-Observation

ICRA 2021poster

The learning-from-observation (LfO) framework aims to map human demonstrations to a robot to reduce programming effort. To this end, an LfO system encodes a human demonstration into a series of execution units for a robot, which are referred to as task models. Although previous research has proposed…

Cited by 17SourceScholar