ICLR 2018workshop6 citations

Not-So-CLEVR: Visual Relations Strain Feedforward Neural Networks

Junkyung Kim, Matthew Ricci, Thomas Serre

Abstract

The robust and efficient recognition of visual relations in images is a hallmark of biological vision. Here, we argue that, despite recent progress in visual recognition, modern machine vision algorithms are severely limited in their ability to learn visual relations. Through controlled experiments, we demonstrate that visual-relation problems strain convolutional neural networks (CNNs). The networks eventually break altogether when rote memorization becomes impossible such as when the intra-class variability exceeds their capacity. We further show that another type of feedforward network, called a relational network (RN), which was shown to successfully solve seemingly difficult visual question answering (VQA) problems on the CLEVR datasets, suffers similar limitations. Motivated by the comparable success of biological vision, we argue that feedback mechanisms including working memory and attention are the key computational components underlying abstract visual reasoning.

Visual RelationsVisual ReasoningSVRTAttentionWorking MemoryConvolutional Neural NetworkDeep LearningRelational Network
BibTeX
@misc{
kim2018notsoclevr,
title={Not-So-{CLEVR}: Visual Relations Strain Feedforward Neural Networks},
author={Junkyung Kim and Matthew Ricci and Thomas Serre},
year={2018},
url={https://openreview.net/forum?id=HymuJz-A-},
}
Not-So-CLEVR: Visual Relations Strain Feedforward Neural Networks · ICLR 2018