← Search

Scott Cohen

44 accepted papers

2026

LightMover: Generative Light Movement with Color and Intensity Controls

CVPR 2026

We present LightMover, a framework for controllable light manipulation in single images that leverages video diffusion priors to produce physically plausible illumination changes without re-rendering the scene. We formulate light editing as a sequence-to-sequence prediction problem in visual token s

Cited by 0SourceScholar
2025

CompleteMe: Reference-based Human Image Completion

ICCV 2025poster

Recent methods for human image completion can reconstruct plausible body shapes but often fail to preserve unique details, such as specific clothing patterns or distinctive accessories, without explicit reference images. Even state-of-the-art reference-based inpainting approaches struggle to accurat…

Cited by 0SourcePDFScholar
2025

MetaShadow: Object-Centered Shadow Detection, Removal, and Synthesis

CVPR 2025poster

Shadows are often underconsidered or even ignored in image editing applications, limiting the realism of the edited results. In this paper, we introduce MetaShadow, a three-in-one versatile framework that enables detection, removal, and controllable synthesis of shadows in natural images in an objec…

Cited by 2SourcePDFScholar
2025

Refine-by-Align: Reference-Guided Artifacts Refinement through Semantic Alignment

ICLR 2025poster

Personalized image generation has emerged from the recent advancements in generative models. However, these generated personalized images often suffer from localized artifacts such as incorrect logos, reducing fidelity and fine-grained identity details of the generated results. Furthermore, there is…

Cited by 1SourcePDFScholar
2025

The Photographer's Eye: Teaching Multimodal Large Language Models to See, and Critique Like Photographers

CVPR 2025poster

Photographer, curator, and former director of photography at the Museum of Modern Art (MoMA), John Szarkowski remarked in *William Eggleston's Guide*, "While editing directly from life, photographers have found it too difficult to see simultaneously both the blue and the sky." Szarkowski insightfull…

Cited by 0SourcePDFScholar
2024

FairDeDup: Detecting and Mitigating Vision-Language Fairness Disparities in Semantic Dataset Deduplication

CVPR 2024poster

Recent dataset deduplication techniques have demonstrated that content-aware dataset pruning can dramatically reduce the cost of training Vision-Language Pretrained (VLP) models without significant performance losses compared to training on the original dataset. These results have been based on prun…

Cited by 6SourcePDFScholar
2024

FineMatch: Aspect-based Fine-grained Image and Text Mismatch Detection and Correction

ECCV 2024poster

"Recent progress in large-scale pre-training has led to the development of advanced vision-language models (VLMs) with remarkable proficiency in comprehending and generating multimodal content. Despite the impressive ability to perform complex reasoning for VLMs, current models often struggle to eff…

2024

IMPRINT: Generative Object Compositing by Learning Identity-Preserving Representation

CVPR 2024poster

Generative object compositing emerges as a promising new avenue for compositional image editing. However the requirement of object identity preservation poses a significant challenge limiting practical usage of most existing methods. In response this paper introduces IMPRINT a novel diffusion-based…

Cited by 29SourcePDFScholar
2023

ObjectStitch: Object Compositing With Diffusion Model

CVPR 2023poster

Object compositing based on 2D images is a challenging problem since it typically involves multiple processing stages such as color harmonization, geometry correction and shadow generation to generate realistic results. Furthermore, annotating training data pairs for compositing requires substantial…

Cited by 94SourcePDFScholar
2023

TopNet: Transformer-Based Object Placement Network for Image Compositing

CVPR 2023poster

We investigate the problem of automatically placing an object into a background image for image compositing. Given a background image and a segmented object, the goal is to train a model to predict plausible placements (location and scale) of the object for compositing. The quality of the composite…

Cited by 18SourcePDFScholar
2022

GALA: Toward Geometry-and-Lighting-Aware Object Search for Compositing

ECCV 2022poster

"Compositing-aware object search aims to find the most compatible objects for compositing given a background image and a query bounding box. Previous works focus on learning compatibility between the foreground object and background, but fail to learn other important factors from large-scale data, i…

Cited by 7SourcePDFScholar
2022

Image Inpainting with Cascaded Modulation GAN and Object-Aware Training

ECCV 2022poster

"Recent image inpainting methods have made great progress but often struggle to generate plausible image structures when dealing with large holes in complex images. This is partially due to the lack of effective network structures that can capture both the long-range dependency and high-level semant…

2022

Improving Closed and Open-Vocabulary Attribute Prediction Using Transformers

ECCV 2022poster

"We study recognizing attributes for objects in visual scenes. We consider attributes to be any phrases that describe an object’s physical and semantic properties, and its relationships with other objects. Existing work studies attribute prediction in a closed setting with a fixed set of attributes,…

Cited by 24SourcePDFScholar
2021

AESOP: Abstract Encoding of Stories, Objects, and Pictures

ICCV 2021poster

Visual storytelling and story comprehension are uniquely human skills that play a central role in how we learn about and experience the world. Despite remarkable progress in recent years in synthesis of visual and textual content in isolation and learning effective joint visual-linguistic representa…

Cited by 19PDFcodeScholar
2021

Learning To Predict Visual Attributes in the Wild

CVPR 2021poster

Visual attributes constitute a large portion of information contained in a scene. Objects can be described using a wide variety of attributes which portray their visual appearance (color, texture), geometry (shape, size, posture), and other intrinsic properties (state, action). Existing work is most…

Cited by 132PDFScholar
2020

PhraseClick: Toward Achieving Flexible Interactive Segmentation by Phrase and Click

ECCV 2020poster

Existing interactive object segmentation methods mainly take spatial interactions such as bounding boxes or clicks as input. However, these interactions do not contain information about explicit attributes of the target-of-interest and thus cannot quickly specify what the selected object exactly is,…

Cited by 64SourcePDFScholar
2020

PhraseCut: Language-Based Image Segmentation in the Wild

CVPR 2020poster

We consider the problem of segmenting image regions given a natural language phrase, and study it on a novel dataset of 77,262 images and 345,486 phrase-region pairs. Our dataset is collected on top of the Visual Genome dataset and uses the existing annotations to generate a challenging set of refer…

Cited by 130PDFcodeScholar
2019

MultiSeg: Semantically Meaningful, Scale-Diverse Segmentations From Minimal User Input

ICCV 2019poster

Existing deep learning-based interactive image segmentation approaches typically assume the target-of-interest is always a single object and fail to account for the potential diversity in user expectations, thus requiring excessive user input when it comes to segmenting an object part or a group of…

Cited by 45PDFcodeScholar
2019

When Color Constancy Goes Wrong: Correcting Improperly White-Balanced Images

CVPR 2019poster

This paper focuses on correcting a camera image that has been improperly white-balanced. This situation occurs when a camera's auto white balance fails or when the wrong manual white-balance setting is used. Even after decades of computational color constancy research, there are no effective solutio…

Cited by 161PDFScholar
2018

Concept Mask: Large-Scale Segmentation from Semantic Concepts

ECCV 2018poster

Existing works on semantic segmentation typically consider a small number of labels, ranging from tens to a few hundreds. With a large number of labels, training and evaluation of such task become extremely challenging due to correlation between labels and lack of datasets with complete annotations.…

Cited by 21SourcePDFScholar
2018

DVQA: Understanding Data Visualizations via Question Answering

CVPR 2018poster

Bar charts are an effective way to convey numeric information, but today's algorithms cannot parse them. Existing methods fail when faced with even minor variations in appearance. Here, we present DVQA, a dataset that tests many aspects of bar chart understanding in a question answering framework. U…

2018

Discriminability Objective for Training Descriptive Captions

CVPR 2018poster

One property that remains lacking in image captions generated by contemporary methods is discriminability: being able to tell two images apart given the caption for one of them. We propose a way to improve this aspect of caption generation. By incorporating into the captioning training objective a l…

2018

Interactive Boundary Prediction for Object Selection

ECCV 2018poster

Interactive image segmentation is critical for many image editing tasks. While recent advanced methods on interactive segmentation focus on the region-based paradigm, more traditional boundary-based methods such as Intelligent Scissor are still popular in practice as they allow users to have active…

Cited by 65SourcePDFScholar
2018

Start, Follow, Read: End-to-End Full-Page Handwriting Recognition

ECCV 2018poster

Despite decades of research, offline handwriting recognition (HWR) of degraded historical documents remains a challenging problem, which if solved could greatly improve the searchability of online cultural heritage archives. HWR models are often limited by the accuracy of the preceding steps of text…

2018

YouTube-VOS: Sequence-to-Sequence Video Object Segmentation

ECCV 2018poster

Learning long-term spatial-temporal features are critical for many video analysis tasks. However, existing video segmentation methods predominantly rely on static image segmentation techniques, and methods capturing temporal dependency for segmentation have to depend on pretrained optical flow model…

Cited by 594SourcePDFScholar
2017

Skeleton Key: Image Captioning by Skeleton-Attribute Decomposition

CVPR 2017poster

Recently, there has been a lot of interest in automatically generating descriptions for an image. Most existing language-model based approaches for this task learn to generate an image description word by word in its original word order. However, for humans, it is more natural to locate the objects…

Cited by 147PDFScholar
2016

Object Contour Detection With a Fully Convolutional Encoder-Decoder Network

CVPR 2016spotlight

We develop a deep learning algorithm for contour detection with a fully convolutional encoder-decoder network. Different from previous low-level edge detection, our algorithm focuses on detecting higher-level object contours. Our network is trained end-to-end on PASCAL VOC with refined ground truth…

Cited by 478PDFScholar
2016

SURGE: Surface Regularized Geometry Estimation from a Single Image

NeurIPS 2016poster

This paper introduces an approach to regularize 2.5D surface normal and depth predictions at each pixel given a single input image. The approach infers and reasons about the underlying 3D planar surfaces depicted in the image to snap predicted normals and depths to inferred planar surfaces, all whil…

Cited by 103SourcePDFScholar
2016

Two Illuminant Estimation and User Correction Preference

CVPR 2016poster

This paper examines the problem of white-balance correction when a scene contains two illuminations. This is a two step process: 1) estimate the two illuminants; and 2) correct the image. Existing methods attempt to estimate a spatially varying illumination map, however, results are error prone a…

Cited by 43PDFScholar
2015

Beyond White: Ground Truth Colors for Color Constancy Correction

ICCV 2015poster

A limitation in color constancy research is the inability to establish ground truth colors for evaluating corrected images. Many existing datasets contain images of scenes with a color chart included; however, only the chart's neutral colors (grayscale patches) are used to provide the ground truth f…

Cited by 64PDFScholar
2015

Effective Learning-Based Illuminant Estimation Using Simple Features

CVPR 2015poster

Illumination estimation is the process of determining the chromaticity of the illumination in an imaged scene in order to remove undesirable color casts through white-balancing. While computational color constancy is a well-studied topic in computer vision, it remains challenging due to the ill-po…

Cited by 179SourcePDFScholar
2015

Joint Object and Part Segmentation Using Deep Learned Potentials

ICCV 2015poster

Segmenting semantic objects from images and parsing them into their respective semantic parts are fundamental steps towards detailed object understanding in computer vision. In this paper, we propose a joint solution that tackles semantic object and part segmentation simultaneously, in which higher…

Cited by 141PDFScholar
2015

PatchCut: Data-Driven Object Segmentation via Local Shape Transfer

CVPR 2015poster

Object segmentation is highly desirable for image understanding and editing. Current interactive tools require a great deal of user effort while automatic methods are usually limited to images of special object categories or with high color contrast. In this paper, we propose a data-driven algorithm…

Cited by 25SourcePDFScholar
2015

Towards Unified Depth and Semantic Prediction From a Single Image

CVPR 2015poster

Depth estimation and semantic segmentation are two fundamental problems in image understanding. While the two tasks are strongly correlated and mutually beneficial, they are usually solved separately or sequentially. Motivated by the complementary properties of the two tasks, we propose a unified fr…