← Search

Gabriela Sejnova

3 accepted papers

2024

Bridging Language, Vision and Action: Multimodal VAEs in Robotic Manipulation Tasks

IROS 2024poster

In this work, we focus on unsupervised vision-language-action mapping in the area of robotic manipulation. Recently, multiple approaches employing pre-trained large language and vision models have been proposed for this task. However, they are computationally demanding and require careful fine-tunin…

Cited by 3SourcecodeScholar
2023

Imitrob: Imitation Learning Dataset for Training and Evaluating 6D Object Pose Estimators

RA-L 2023

This letter introduces a dataset for training and evaluating methods for 6D pose estimation of hand-held tools in task demonstrations captured by a standard RGB camera. Despite the significant progress of 6D pose estimation methods, their performance is usually limited for heavily occluded objects,

Cited by 7SourcecodeScholar
2019

Exploring logical consistency and viewport sensitivity in compositional VQA models

IROS 2019poster

The most recent architectures for Visual Question Answering (VQA), such as TbD or DDRprog, have already outperformed human-level accuracy on benchmark datasets (e.g. CLEVR). We administered an advanced analysis of their performance based on novel metrics called consistency (sum of all object feature…

Cited by 2SourceScholar