2017
Unsupervised Linking of Visual Features to Textual Descriptions in Long Manipulation Activities
RA-L 2017
We present a novel unsupervised framework, which links continuous visual features and symbolic textual descriptions of manipulation activity videos. First, we extract the semantic representation of visually observed manipulations by applying a bottom-up approach to the continuous image streams. We t