← Search

Tamara L. Berg

12 accepted papers

2022

End-to-End Visual Editing with a Generatively Pre-trained Artist

ECCV 2022poster

"We consider the targeted image editing problem, namely blending a region in a source image with a driver image that specifies the desired change. Differently from prior works, we solve this problem by learning a conditional probability distribution of the edits, end-to-end in code space. Training s…

Cited by 6SourcePDFScholar
2021

Connecting What To Say With Where To Look by Modeling Human Attention Traces

CVPR 2021poster

We introduce a unified framework to jointly model images, text, and human attention traces. Our work is built on top of the recent Localized Narratives annotation framework, where each word of a given caption is paired with a mouse trace segment. We propose two novel tasks: (1) predict a trace given…

Cited by 32PDFcodeScholar
2021

Less Is More: ClipBERT for Video-and-Language Learning via Sparse Sampling

CVPR 2021poster

The canonical approach to video-and-language learning (e.g., video question answering) dictates a neural model to learn from offline-extracted dense video features from vision models and text features from language models. These feature extractors are trained independently and usually on tasks diffe…

Cited by 771PDFcodeScholar
2020

TVR: A Large-Scale Dataset for Video-Subtitle Moment Retrieval

ECCV 2020poster

We introduce TV show Retrieval (TVR), a new multimodal retrieval dataset. TVR requires systems to understand both videos and their associated subtitle (dialogue) texts, making it more realistic. The dataset contains 109K queries collected on 21.8K videos from 6 TV shows of diverse genres, where each…

Cited by 340SourcePDFScholar
2019

IMP: Instance Mask Projection for High Accuracy Semantic Segmentation of Things

ICCV 2019poster

In this work, we present a new operator, called Instance Mask Projection (IMP), which projects a predicted instance segmentation as a new feature for semantic segmentation. It also supports back propagation and is trainable end-to end. By adding this operator, we introduce a new way to combine top-d…

Cited by 22PDFScholar
2019

Multi-Target Embodied Question Answering

CVPR 2019poster

Embodied Question Answering (EQA) is a relatively new task where an agent is asked to answer questions about its environment from egocentric perception. EQA as introduced in [8] makes the fundamental assumption that every question, e.g., "what color is the car?", has exactly one target ("car") bein…

Cited by 130PDFcodeScholar
2018

MAttNet: Modular Attention Network for Referring Expression Comprehension

CVPR 2018poster

In this paper, we address referring expression comprehension: localizing an image region described by a natural language expression. While most recent work treats expressions as a single unit, we propose to decompose them into three modular components related to subject appearance, location, and re…

2018

Visual to Sound: Generating Natural Sound for Videos in the Wild

CVPR 2018poster

As two of the five traditional human senses (sight, hearing, taste, smell, and touch), vision and sound are basic sources through which humans understand the world. Often correlated during natural events, these two modalities combine to jointly affect human perception. In this paper, we pose the tas…

Cited by 261SourcePDFScholar
2015

Visual Madlibs: Fill in the Blank Description Generation and Question Answering

ICCV 2015poster

In this paper, we introduce a new dataset consisting of 360,001 focused natural language descriptions for 10,738 images. This dataset, the Visual Madlibs dataset, is collected using automatically produced fill-in-the-blank templates designed to gather targeted descriptions about: people and objects…

Cited by 181PDFcodeScholar
2015

Where to Buy It: Matching Street Clothing Photos in Online Shops

ICCV 2015oral

In this paper, we define a new task, Exact Street to Shop, where our goal is to match a real-world example of a garment item to the same item in an online shop. This is an extremely challenging task due to visual differences between street photos (pictures of people wearing clothing in everyday unco…

Cited by 578PDFScholar