← Search

Hsiao-Yu Fish Tung

4 accepted papers

2020

Embodied Language Grounding With 3D Visual Feature Representations

CVPR 2020poster

We propose associating language utterances to 3D visual abstractions of the scene they describe. The 3D visual abstractions are encoded as 3-dimensional visual feature maps. We infer these 3D visual scene feature maps from RGB images of the scene via view prediction: when the generated 3D scene feat…

Cited by 23PDFScholar
2020

Learning from Unlabelled Videos Using Contrastive Predictive Neural 3D Mapping

ICLR 2020poster

Predictive coding theories suggest that the brain learns by predicting observations at various levels of abstraction. One of the most basic prediction tasks is view prediction: how would a given scene look from an alternative viewpoint? Humans excel at this task. Our ability to imagine and fill in m…

Cited by 30SourcecodeScholar
2017

Adversarial Inverse Graphics Networks: Learning 2D-To-3D Lifting and Image-To-Image Translation From Unpaired Supervision

ICCV 2017poster

Researchers have developed excellent feed-forward models that learn to map images to desired outputs, such as to the images' latent factors, or to other images, using supervised learning. Learning such mappings from unlabelled data, or improving upon supervised models by exploiting unlabelled data,…

Cited by 170PDFScholar