← Search

Andrew Gilbert

7 accepted papers

2025

Multitwine: Multi-Object Compositing with Text and Layout Control

CVPR 2025highlight

We introduce the first generative model capable of simultaneous multi-object compositing, guided by both text and layout. Our model allows for the addition of multiple objects within a scene, capturing a range of interactions from simple positional relations (e.g., next to, in front of) to complex…

Cited by 1SourcePDFScholar
2024

Thinking Outside the BBox: Unconstrained Generative Object Compositing

ECCV 2024poster

"Compositing an object into an image involves multiple non-trivial sub-tasks such as object placement and scaling, color/lighting harmonization, viewpoint/geometry adjustment, and shadow/reflection generation. Recent generative image compositing methods leverage diffusion models to handle multiple s…

Cited by 9SourcePDFScholar
2022

StyleBabel: Artistic Style Tagging and Captioning

ECCV 2022poster

"We present StyleBabel, a unique open access dataset of natural language captions and free-form tags describing the artistic style of over 135K digital artworks, collected via a novel participatory method from experts studying at specialist art and design schools. StyleBabel was collected via an ite…

Cited by 15SourcePDFScholar
2021

ALADIN: All Layer Adaptive Instance Normalization for Fine-Grained Style Similarity

ICCV 2021poster

We present ALADIN (All Layer AdaIN); a novel architecture for searching images based on the similarity of their artistic style. Representation learning is critical to visual search, where distance in the learned search embedding reflects image similarity. Learning an embedding that discriminates fin…

Cited by 32PDFScholar
2018

Deep Autoencoder for Combined Human Pose Estimation and Body Model Upscaling

ECCV 2018poster

We present a method for simultaneously estimating 3D human pose and body shape from a sparse set of wide-baseline camera views. We train a symmetric convolutional autoencoder with a dual loss that enforces learning of a latent representation that encodes skeletal joint positions, and at the same tim…

Cited by 75SourcePDFScholar
2018

Disentangling Structure and Aesthetics for Style-Aware Image Completion

CVPR 2018poster

Content-aware image completion or in-painting is a fundamental tool for the correction of defects or removal of objects in images. We propose a non-parametric in-painting algorithm that enforces both structural and aesthetic (style) consistency within the resulting image. Our contributions are two…

Cited by 15SourcePDFScholar
2018

Volumetric performance capture from minimal camera viewpoints

ECCV 2018poster

We present a convolutional autoencoder that enables high fidelity volumetric reconstructions of human performance to be captured from multi-view video comprising only a small set of camera views. Our method yields similar end-to-end reconstruction error to that of a probabilistic visual hull compute…

Cited by 68SourcePDFScholar