← Search

Sarah Parisot

12 accepted papers

2024

Generating compositional scenes via Text-to-image RGBA Instance Generation

NeurIPS 2024poster

Text-to-image diffusion generative models can generate high quality images at the cost of tedious prompt engineering. Controllability can be improved by introducing layout conditioning, however existing methods lack layout editing ability and fine-grained control over object attributes. The concept…

Cited by 0SourcePDFScholar
2024

MULAN: A Multi Layer Annotated Dataset for Controllable Text-to-Image Generation

CVPR 2024poster

Text-to-image generation has achieved astonishing results yet precise spatial controllability and prompt fidelity remain highly challenging. This limitation is typically addressed through cumbersome prompt engineering scene layout conditioning or image editing techniques which often require hand dra…

2022

Long-Tail Recognition via Compositional Knowledge Transfer

CVPR 2022poster

In this work, we introduce a novel strategy for long-tail recognition that addresses the tail classes' few-shot problem via training-free knowledge transfer. Our objective is to transfer knowledge acquired from information-rich common classes to semantically similar, and yet data-hungry, rare classe…

Cited by 39PDFScholar
2020

A Multi-Hypothesis Approach to Color Constancy

CVPR 2020poster

Contemporary approaches frame the color constancy problem as learning camera specific illuminant mappings. While high accuracy can be achieved on camera specific data, these models depend on camera spectral sensitivity and typically exhibit poor generalisation to new devices. Additionally, regressio…

Cited by 65PDFScholar
2020

DeepLPF: Deep Local Parametric Filters for Image Enhancement

CVPR 2020poster

Digital artists often improve the aesthetic quality of digital photographs through manual retouching. Beyond global adjustments, professional image editing programs provide local adjustment tools operating on specific parts of an image. Options include parametric (graduated, radial filters) and unco…

Cited by 289PDFScholar
2020

Few-Shot Single-View 3-D Object Reconstruction with Compositional Priors

ECCV 2020poster

The impressive performance of deep convolutional neural networks in single-view 3D reconstruction suggests that these models perform non-trivial reasoning about the 3D structure of the output space. However, recent work has challenged this belief, showing that complex encoder-decoder architectures p…

Cited by 27SourcePDFScholar
2020

Low Light Video Enhancement using Synthetic Data Produced with an Intermediate Domain Mapping

ECCV 2020poster

Advances in low-light video RAW-to-RGB translation are opening up the possibility of fast low-light imaging on commodity devices (e.g. smartphone cameras) without the need for a tripod. However,it is challenging to collect the required paired short-long exposure frames to learn a supervised mapping.…

2020

Many-shot from Low-shot: Learning to Annotate using Mixed Supervision for Object Detection

ECCV 2020poster

Object detection has witnessed significant progress by relying on large, manually annotated datasets. Annotating such datasets is highly time consuming and expensive, which motivates the development of weakly supervised and few-shot object detection methods. However, these methods largely underperfo…

Cited by 18SourcePDFScholar
2020

More Classifiers, Less Forgetting: A Generic Multi-classifier Paradigm for Incremental Learning

ECCV 2020poster

Less Forgetting: A Generic Multi-classifier Paradigm for Incremental Learning","Overcoming catastrophic forgetting in neural networks is a long-standing and core research objective for incremental learning. Notable studies have shown regularization strategies enable the network to remember previousl…

2020

Unsupervised Model Personalization While Preserving Privacy and Scalability: An Open Problem

CVPR 2020poster

This work investigates the task of unsupervised model personalization, adapted to continually evolving, unlabeled local user images. We consider the practical scenario where a high capacity server interacts with a myriad of resource-limited edge devices, imposing strong requirements on scalability a…

Cited by 35PDFcodeScholar
2018

Learning Conditioned Graph Structures for Interpretable Visual Question Answering

NeurIPS 2018poster

Visual Question answering is a challenging problem requiring a combination of concepts from Computer Vision and Natural Language Processing. Most existing approaches use a two streams strategy, computing image and question features that are consequently merged using a variety of techniques. Nonethel…