← Search

Atsushi Kanehira

7 accepted papers

2024

GPT-4V(ision) for Robotics: Multimodal Task Planning From Human Demonstration

RA-L 2024

We introduce a pipeline that enhances a general-purpose Vision Language Model, GPT-4V(ision), to facilitate one-shot visual teaching for robotic manipulation. This system analyzes videos of humans performing tasks and outputs executable robot programs that incorporate insights into affordances. The

Cited by 117SourceScholar
2021

Hierarchical Lovasz Embeddings for Proposal-Free Panoptic Segmentation

CVPR 2021poster

Panoptic segmentation brings together two separate tasks: instance and semantic segmentation. Although they are related, unifying them faces an apparent paradox: how to learn simultaneously instance-specific and category-specific (i.e. instance-agnostic) representations jointly. Hence, state-of-the-…

Cited by 10PDFScholar
2019

Multimodal Explanations by Predicting Counterfactuality in Videos

CVPR 2019oral

This study addresses generating counterfactual explanations with multimodal information. Our goal is not only to classify a video into a specific category, but also to provide explanations on why it is not categorized to a specific class with combinations of visual-linguistic information. Requiremen…

Cited by 48PDFScholar
2016

Recognizing Activities of Daily Living With a Wrist-Mounted Camera

CVPR 2016spotlight

We present a novel dataset and a novel algorithm for recognizing activities of daily living (ADL) from a first-person wearable camera. Handled objects are crucially important for egocentric ADL recognition. For specific examination of objects related to users' actions separately from other objects i…

Cited by 68PDFScholar