← Search

Ahmed Elgammal

16 accepted papers

2022

Proxy Learning of Visual Concepts of Fine Art Paintings from Styles through Language Models

AAAI 2022technical

We present a machine learning system that can quantify fine art paintings with a set of visual elements and principles of art. The formal analysis is fundamental for understanding art, but developing such a system is challenging. Paintings have high visual complexities, but it is also difficult to c…

Cited by 2SourcePDFScholar
2021

TIME: Text and Image Mutual-Translation Adversarial Networks

AAAI 2021technical

Focusing on text-to-image (T2I) generation, we propose Text and Image Mutual-Translation Adversarial Networks (TIME), a lightweight but effective model that jointly learns a T2I generator G and an image captioning discriminator D under the Generative Adversarial Network framework. While previous met…

Cited by 39SourcePDFScholar
2021

Towards Faster and Stabilized GAN Training for High-fidelity Few-shot Image Synthesis

ICLR 2021poster

Training Generative Adversarial Networks (GAN) on high-fidelity images usually requires large-scale GPU-clusters and a vast number of training images. In this paper, we study the few-shot image synthesis task for GAN with minimum computing cost. We propose a light-weight GAN structure that gains sup…

2019

Graphical Contrastive Losses for Scene Graph Parsing

CVPR 2019poster

Most scene graph parsers use a two-stage pipeline to detect visual relationships: the first stage detects entities, and the second predicts the predicate for each entity pair using a softmax distribution. We find that such pipelines, trained with only a cross entropy loss over predicate classes, suf…

Cited by 289PDFScholar
2019

Learning Feature-to-Feature Translator by Alternating Back-Propagation for Generative Zero-Shot Learning

ICCV 2019poster

We investigate learning feature-to-feature translator networks by alternating back-propagation as a general-purpose solution to zero-shot learning (ZSL) problems. It is a generative model-based ZSL framework. In contrast to models based on generative adversarial networks (GAN) or variational autoenc…

Cited by 127PDFcodeScholar
2019

Semantic-Guided Multi-Attention Localization for Zero-Shot Learning

NeurIPS 2019poster

Zero-shot learning extends the conventional object classification to the unseen class recognition by introducing semantic representations of classes. Existing approaches predominantly focus on learning the proper mapping function for visual-semantic embedding, while neglecting the effect of learning…

Cited by 180SourcePDFScholar
2018

A Generative Adversarial Approach for Zero-Shot Learning From Noisy Texts

CVPR 2018poster

Most existing zero-shot learning methods consider the problem as a visual semantic embedding one. Given the demonstrated capability of Generative Adversarial Networks(GANs) to generate images, we instead leverage GANs to imagine unseen categories from text descriptions and hence recognize novel clas…

Cited by 499SourcePDFScholar
2018

Disconnected Manifold Learning for Generative Adversarial Networks

NeurIPS 2018poster

Natural images may lie on a union of disjoint manifolds rather than one globally connected manifold, and this can cause several difficulties for the training of common Generative Adversarial Networks (GANs). In this work, we first show that single generator GANs are unable to correctly model a distr…

Cited by 116SourcePDFScholar
2017

Link the Head to the "Beak": Zero Shot Learning From Noisy Text Description at Part Precision

CVPR 2017poster

In this paper, we study learning visual classifiers from unstructured text description at part precision with no training images. We show that visual text terms can be encouraged to attend to its relevant parts, while image connections to non-visual text terms vanishes without any supervision. Thi…

Cited by 158PDFScholar
2016

A Comparative Analysis and Study of Multiview CNN Models for Joint Object Categorization and Pose Estimation

ICML 2016poster

In the Object Recognition task, there exists a dichotomy between the categorization of objects and estimating object pose, where the former necessitates a view-invariant representation, while the latter requires a representation capable of capturing pose information over different categories of obje…

Cited by 44SourcePDFScholar
2016

SPDA-CNN: Unifying Semantic Part Detection and Abstraction for Fine-Grained Recognition

CVPR 2016poster

Most convolutional neural networks (CNNs) lack midlevel layers that model semantic parts of objects. This limits CNN-based methods from reaching their full potential in detecting and utilizing small semantic parts in recognition. Introducing such mid-level layers can facilitate the extraction of par…

Cited by 382PDFScholar