Bridging Language, Vision and Action: Multimodal VAEs in Robotic Manipulation Tasks
In this work, we focus on unsupervised vision-language-action mapping in the area of robotic manipulation. Recently, multiple approaches employing pre-trained large language and vision models have been proposed for this task. However, they are computationally demanding and require careful fine-tunin…