← Search

Michal Vavrecka

2 accepted papers

2024

Bridging Language, Vision and Action: Multimodal VAEs in Robotic Manipulation Tasks

IROS 2024poster

In this work, we focus on unsupervised vision-language-action mapping in the area of robotic manipulation. Recently, multiple approaches employing pre-trained large language and vision models have been proposed for this task. However, they are computationally demanding and require careful fine-tunin…

Cited by 3SourcecodeScholar
2019

Exploring logical consistency and viewport sensitivity in compositional VQA models

IROS 2019poster

The most recent architectures for Visual Question Answering (VQA), such as TbD or DDRprog, have already outperformed human-level accuracy on benchmark datasets (e.g. CLEVR). We administered an advanced analysis of their performance based on novel metrics called consistency (sum of all object feature…

Cited by 2SourceScholar