ICLR 2017workshop73 citations

Generating Interpretable Images with Controllable Structure

Scott Reed, Aäron van den Oord, Nal Kalchbrenner, Victor Bapst, Matt Botvinick, Nando de Freitas

Abstract

We demonstrate improved text-to-image synthesis with controllable object locations using an extension of Pixel Convolutional Neural Networks (PixelCNN). In addition to conditioning on text, we show how the model can generate images conditioned on part keypoints and segmentation masks. The character-level text encoder and image generation network are jointly trained end-to-end via maximum likelihood. We establish quantitative baselines in terms of text and structure-conditional pixel log-likelihood for three data sets: Caltech-UCSD Birds (CUB), MPII Human Pose (MHP), and Common Objects in Context (MS-COCO).

Deep learningComputer visionMulti-modal learningNatural language processing
BibTeX
@misc{
lee2017making,
title={Making Stochastic Neural Networks from Deterministic Ones},
author={Kimin Lee and Jaehyung Kim and Song Chong and Jinwoo Shin},
year={2017},
url={https://openreview.net/forum?id=B1akgy9xx}
}
Generating Interpretable Images with Controllable Structure · ICLR 2017