← Search

Alexander C. Berg

17 accepted papers

2026

Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens

CVPR 2026

Current text-to-image models struggle to provide precise camera control using natural language alone. In this work, we present a framework for precise camera control with global scene understanding in text-to-image generation by learning parametric camera tokens. We fine-tune image generation models

Cited by 0SourcecodeScholar
2024

Improved Visual Grounding through Self-Consistent Explanations

CVPR 2024poster

Vision-and-language models trained to match images with text can be combined with visual explanation methods to point to the locations of specific objects in an image. Our work shows that the localization --"grounding'"-- abilities of these models can be further improved by finetuning for self-consi…

Cited by 15SourcePDFScholar
2022

Point-Level Region Contrast for Object Detection Pre-Training

CVPR 2022oral

In this work we present point-level region contrast, a self-supervised pre-training approach for the task of object detection. This approach is motivated by the two key factors in detection: localization and recognition. While accurate localization favors models that operate at the pixel- or point-l…

Cited by 64PDFcodeScholar
2022

Similarity Search for Efficient Active Learning and Search of Rare Concepts

AAAI 2022technical

Many active learning and search approaches are intractable for large-scale industrial settings with billions of unlabeled examples. Existing approaches search globally for the optimal examples to label, scaling linearly or even quadratically with the unlabeled data. In this paper, we improve the com…

Cited by 41SourcePDFScholar
2021

Boundary IoU: Improving Object-Centric Image Segmentation Evaluation

CVPR 2021poster

We present Boundary IoU (Intersection-over-Union), a new segmentation evaluation measure focused on boundary quality. We perform an extensive analysis across different error types and object sizes and show that Boundary IoU is significantly more sensitive than the standard Mask IoU measure to bounda…

Cited by 405PDFcodeScholar
2021

Neural Pseudo-Label Optimism for the Bank Loan Problem

NeurIPS 2021poster

We study a class of classification problems best exemplified by the \emph{bank loan} problem, where a lender decides whether or not to issue a loan. The lender only observes whether a customer will repay a loan if the loan is issued to begin with, and thus modeled decisions affect what data is avail…

Cited by 8SourcePDFScholar
2021

Worldsheet: Wrapping the World in a 3D Sheet for View Synthesis From a Single Image

ICCV 2021poster

We present Worldsheet, a method for novel view synthesis using just a single RGB image as input. The main insight is that simply shrink-wrapping a planar mesh sheet onto the input image, consistent with the learned intermediate depth, captures underlying geometry sufficient to generate photorealisti…

Cited by 85PDFcodeScholar
2019

IMP: Instance Mask Projection for High Accuracy Semantic Segmentation of Things

ICCV 2019poster

In this work, we present a new operator, called Instance Mask Projection (IMP), which projects a predicted instance segmentation as a new feature for semantic segmentation. It also supports back propagation and is trainable end-to end. By adding this operator, we introduce a new way to combine top-d…

Cited by 22PDFScholar
2019

Leveraging Long-Range Temporal Relationships Between Proposals for Video Object Detection

ICCV 2019poster

Single-frame object detectors perform well on videos sometimes, even without temporal context. However, challenges such as occlusion, motion blur, and rare poses of objects are hard to resolve without temporal awareness. Thus, there is a strong need to improve video object detection by considering l…

Cited by 119PDFScholar
2018

Meta-Tracker: Fast and Robust Online Adaptation for Visual Object Trackers

ECCV 2018poster

This paper improves state-of-the-art visual object trackers that use online adaptation. Our core contribution is an offline meta-learning-based method to adjust the initial deep networks used in online adaptation-based tracking. The meta learning is driven by the goal of deep networks that can quick…

2017

A dataset for developing and benchmarking active vision

ICRA 2017poster

We present a new public dataset with a focus on simulating robotic vision tasks in everyday indoor environments using real imagery. The dataset includes 20,000+ RGB-D images and 50,000+ 2D bounding boxes of object instances densely captured in 9 unique scenes. We train a fast object category detecto…

Cited by 240SourceScholar
2017

Transformation-Grounded Image Generation Network for Novel 3D View Synthesis

CVPR 2017poster

We present a transformation-grounded image generation network for novel 3D view synthesis from a single image. Our approach first explicitly infers the parts of the geometry visible both in the input and novel views and then casts the remaining synthesis problem as image completion. Specifically, we…

Cited by 346PDFcodeScholar
2015

MatchNet: Unifying Feature and Metric Learning for Patch-Based Matching

CVPR 2015poster

Motivated by recent successes on learning feature representations and on learning feature comparison functions, we propose a unified approach to combining both for training a patch matching system. Our system, dubbed MatchNet, consists of a deep convolutional network that extracts features from pa…

2015

PAIGE: PAirwise Image Geometry Encoding for Improved Efficiency in Structure-From-Motion

CVPR 2015poster

Large-scale Structure-from-Motion systems typically spend major computational effort on pairwise image matching and geometric verification in order to discover connected components in large-scale, unordered image collections. In recent years, the research community has spent significant effort on im…

Cited by 44SourcePDFScholar
2015

Visual Madlibs: Fill in the Blank Description Generation and Question Answering

ICCV 2015poster

In this paper, we introduce a new dataset consisting of 360,001 focused natural language descriptions for 10,738 images. This dataset, the Visual Madlibs dataset, is collected using automatically produced fill-in-the-blank templates designed to gather targeted descriptions about: people and objects…

Cited by 181PDFcodeScholar
2015

Where to Buy It: Matching Street Clothing Photos in Online Shops

ICCV 2015oral

In this paper, we define a new task, Exact Street to Shop, where our goal is to match a real-world example of a garment item to the same item in an online shop. This is an extremely challenging task due to visual differences between street photos (pictures of people wearing clothing in everyday unco…

Cited by 578PDFScholar