← Search

Anelia Angelova

40 accepted papers

2026

CURVE: A Benchmark for Cultural and Multilingual Long Video Reasoning

CVPR 2026

Recent advancements in video models have shown tremendous progress, particularly in long video understanding. However, current benchmarks predominantly feature western-centric data and English as the dominant language, introducing significant biases in evaluation. To address this, we introduce CURVE

Cited by 0SourceScholar
2025

VideoComp: Advancing Fine-Grained Compositional and Temporal Alignment in Video-Text Models

CVPR 2025poster

We introduce VideoComp, a benchmark and learning framework for advancing video-text compositionality understanding, aimed at improving vision-language models (VLMs) in fine-grained temporal alignment. Unlike existing benchmarks focused on static image-text compositionality or isolated single-event v…

2024

3D Open-Vocabulary Panoptic Segmentation with 2D-3D Vision-Language Distillation

ECCV 2024poster

"3D panoptic segmentation is a challenging perception task, especially in autonomous driving. It aims to predict both semantic and instance annotations for 3D points in a scene. Although prior 3D panoptic segmentation approaches have achieved great performance on closed-set benchmarks, generalizing…

Cited by 3SourcePDFScholar
2024

Mirasol3B: A Multimodal Autoregressive Model for Time-Aligned and Contextual Modalities

CVPR 2024poster

One of the main challenges of multimodal learning is the need to combine heterogeneous modalities (e.g. video audio text). For example video and audio are obtained at much higher rates than text and are roughly aligned in time. They are often not synchronized with text which comes as a global contex…

Cited by 23SourcePDFScholar
2024

On Scaling Up a Multilingual Vision and Language Model

CVPR 2024poster

We explore the boundaries of scaling up a multilingual vision and language model both in terms of size of the components and the breadth of its training task mixture. Our model achieves new levels of performance on a wide-range of varied and complex tasks including multiple image-based captioning an…

Cited by 8SourcePDFScholar
2024

Region-centric Image-Language Pretraining for Open-Vocabulary Detection

ECCV 2024poster

"We present a new open-vocabulary detection approach based on region-centric image-language pretraining to bridge the gap between image-level pretraining and open-vocabulary object detection. At the pretraining phase, we incorporate the detector architecture on top of the classification backbone, wh…

2023

Open-Vocabulary Object Detection upon Frozen Vision and Language Models

ICLR 2023poster

We present F-VLM, a simple open-vocabulary object detection method built uponFrozenVision andLanguageModels. F-VLM simplifies the current multi-stage training pipeline by eliminating the need for knowledge distillation or detection-tailored pretraining. Surprisingly, we observe that a frozen VLM:…

Cited by 228SourcePDFScholar
2023

PaLI: A Jointly-Scaled Multilingual Language-Image Model

ICLR 2023top-5%

Effective scaling and a flexible task interface enable large language models to excel at many tasks. We present PaLI, a model that extends this approach to the joint modeling of language and vision. PaLI generates text based on visual and textual inputs, and with this interface performs many vision,…

2023

Region-Aware Pretraining for Open-Vocabulary Object Detection With Vision Transformers

CVPR 2023highlight

We present Region-aware Open-vocabulary Vision Transformers (RO-ViT) -- a contrastive image-text pretraining recipe to bridge the gap between image-level pretraining and open-vocabulary object detection. At the pretraining phase, we propose to randomly crop and resize regions of positional embedding…

Cited by 84SourcePDFScholar
2023

Rethinking Video ViTs: Sparse Video Tubes for Joint Image and Video Learning

CVPR 2023poster

We present a simple approach which can turn a ViT encoder into an efficient video model, which can seamlessly work with both image and video inputs. By sparsely sampling the inputs, the model is able to do training and inference from both inputs. The model is easily scalable and can be adapted to la…

Cited by 81SourcePDFScholar
2022

FindIt: Generalized Localization with Natural Language Queries

ECCV 2022poster

"We propose FindIt, a simple and versatile framework that unifies a variety of visual grounding and localization tasks including referring expression comprehension, text-based localization, and object detection. Key to our architecture is an efficient multi-scale fusion module that unifies the dispa…

2022

Learning Open-World Object Proposals Without Learning to Classify

RA-L 2022

Object proposals have become an integral pre-processing step of many vision pipelines including object detection, weakly supervised detection, object discovery, tracking, etc. Compared to the learning-free methods, learning-based proposals have become popular recently due to the growing interest in

Cited by 158SourcecodeScholar
2022

Mechanical Search on Shelves using a Novel “Bluction” Tool

ICRA 2022poster

Shelves are common in homes, warehouses, and commercial settings due to their storage efficiency. However, this efficiency comes at the cost of reduced visibility and accessibility. When looking from a side (lateral) view of a shelf, most objects will be fully occluded, resulting in a constrained la…

Cited by 24SourceScholar
2022

Video Question Answering with Iterative Video-Text Co-Tokenization

ECCV 2022poster

"Video question answering is a challenging task that requires understanding jointly the language input, the visual information in individual video frames, as well as the temporal information about the events occurring in the video. In this paper, we propose a novel multi-stream video encoder for vid…

Cited by 24SourcePDFScholar
2021

Mechanical Search on Shelves using Lateral Access X-RAY

IROS 2021poster

Finding an occluded object in a lateral access environment such as a shelf or cabinet is a problem that arises in many contexts such as warehouses, retail, healthcare, shipping, and homes. While this problem, known as mechanical search, is well-studied in overhead access environments, lateral access…

Cited by 32SourceScholar
2021

Patch2CAD: Patchwise Embedding Learning for In-the-Wild Shape Retrieval From a Single Image

ICCV 2021poster

3D perception of object shapes from RGB image input is fundamental towards semantic scene understanding, grounding image-based perception in our spatially 3-dimensional real-world environments. To achieve a mapping between image views of objects and 3D shapes, we leverage CAD model priors from exist…

Cited by 36PDFScholar
2021

SMURF: Self-Teaching Multi-Frame Unsupervised RAFT With Full-Image Warping

CVPR 2021poster

We present SMURF, a method for unsupervised learning of optical flow that improves state of the art on all benchmarks by 36% to 40% and even outperforms several supervised approaches such as PWC-Net and FlowNet2. Our method integrates architecture improvements from supervised optical flow, i.e. the…

Cited by 101PDFcodeScholar
2021

TokenLearner: Adaptive Space-Time Tokenization for Videos

NeurIPS 2021poster

In this paper, we introduce a novel visual representation learning which relies on a handful of adaptively learned tokens, and which is applicable to both image and video understanding tasks. Instead of relying on hand-designed splitting strategies to obtain visual tokens and processing a large numb…

Cited by 179SourcePDFScholar
2021

Visionary: Vision architecture discovery for robot learning

ICRA 2021poster

We propose a vision-based architecture search algorithm for robot manipulation learning, which discovers interactions between low dimension action inputs and high dimensional visual inputs. Our approach automatically designs architectures while training on the task – discovering novel ways of combin…

Cited by 12SourceScholar
2020

Adversarial Generative Grammars for Human Activity Prediction

ECCV 2020poster

In this paper we propose an adversarial generative grammar model for future prediction. The objective is to learn a model that explicitly captures temporal dependencies, providing a capability to forecast multiple, distinct future activities. Our adversarial grammar is designed so that it can learn…

Cited by 36SourcePDFScholar
2020

AssembleNet++: Assembling Modality Representations via Attention Connections - Supplementary Material -

ECCV 2020poster

We create a family of powerful video models which are able to: (i) learn interactions between semantic object information and raw appearance and motion features, and (ii) deploy attention in order to better learn the importance of features at each convolutional block of the network. A new network co…

Cited by 1SourcePDFScholar
2020

AssembleNet: Searching for Multi-Stream Neural Connectivity in Video Architectures

ICLR 2020poster

Learning to represent videos is a very challenging task both algorithmically and computationally. Standard video CNN architectures have been designed by directly extending architectures devised for image understanding to include the time dimension, using modules such as 3D convolutions, or by using…

Cited by 121SourcecodeScholar
2020

AttentionNAS: Spatiotemporal Attention Cell Search for Video Classification

ECCV 2020poster

Convolutional operations have two limitations: (1) do not explicitly model where to focus as the same filter is applied to all the positions, and (2) are unsuitable for modeling long-range dependencies as they only operate on a small neighborhood. While both limitations can be alleviated by attentio…

Cited by 56SourcePDFScholar
2020

Differentiable Mapping Networks: Learning Structured Map Representations for Sparse Visual Localization

ICRA 2020poster

Mapping and localization, preferably from a small number of observations, are fundamental tasks in robotics. We address these tasks by combining spatial structure (differentiable mapping) and end-to-end learning in a novel neural network architecture: the Differentiable Mapping Network (DMN). The DM…

Cited by 13SourceScholar
2020

KeyPose: Multi-View 3D Labeling and Keypoint Estimation for Transparent Objects

CVPR 2020poster

Estimating the 3D pose of desktop objects is crucial for applications such as robotic manipulation. Many existing approaches to this problem require a depth map of the object for both training and prediction, which restricts them to opaque, lambertian objects that produce good returns in an RGBD sen…

Cited by 137PDFcodeScholar
2020

Mask2CAD: 3D Shape Prediction by Learning to Segment and Retrieve

ECCV 2020poster

Object recognition has seen significant progress in the image domain, with focus primarily on 2D perception. We propose to leverage existing large-scale datasets of 3D models to understand the underlying 3D structure of objects seen in an image by constructing a CAD-based representation of the objec…

Cited by 95SourcePDFScholar
2020

Unsupervised Monocular Depth Learning in Dynamic Scenes

CoRL 2020

We present a method for jointly training the estimation of depth, ego-motion, and a dense 3D translation field of objects relative to the scene, with monocular photometric consistency being the sole source of supervision. We show that this apparently heavily underdetermined problem can be regularize

2020

What Matters in Unsupervised Optical Flow

ECCV 2020poster

We systematically compare and analyze a set of key components in unsupervised optical flow to identify which photometric loss, occlusion handling, and smoothness regularization is most effective. Alongside this investigation we construct a number of novel improvements to unsupervised flow models, su…

2020

X-Ray: Mechanical Search for an Occluded Object by Minimizing Support of Learned Occupancy Distributions

IROS 2020poster

For applications in e-commerce, warehouses, healthcare, and home service, robots are often required to search through heaps of objects to grasp a specific target object. For mechanical search, we introduce X-Ray, an algorithm based on learned occupancy distributions. We train a neural network using…

Cited by 46SourceScholar
2019

Depth From Videos in the Wild: Unsupervised Monocular Depth Learning From Unknown Cameras

ICCV 2019poster

We present a novel method for simultaneous learning of depth, egomotion, object motion, and camera intrinsics from monocular videos, using only consistency across neighboring video frames as supervision signal. Similarly to prior work, our method learns by applying differentiable warping to frames a…

Cited by 483PDFcodeScholar
2019

ShapeMask: Learning to Segment Novel Objects by Refining Shape Priors

ICCV 2019oral

Instance segmentation aims to detect and segment individual objects in a scene. Most existing methods rely on precise mask annotations of every category. However, it is difficult and costly to segment objects in novel categories because a large number of mask annotations is required. We introduce Sh…

Cited by 156PDFcodeScholar
2018

Unsupervised Learning of Depth and Ego-Motion From Monocular Video Using 3D Geometric Constraints

CVPR 2018poster

We present a novel approach for unsupervised learning of depth and ego-motion from monocular video. Unsupervised learning removes the need for separate supervisory signals (depth or ego-motion ground truth, or multi-view video). Prior work in unsupervised depth learning uses pixel-wise or gradient-…

2017

Deep Value Networks Learn to Evaluate and Iteratively Refine Structured Outputs

ICML 2017poster

We approach structured output prediction by optimizing a deep value network (DVN) to precisely estimate the task loss on different output configurations for a given input. Once the model is trained, we perform inference by gradient descent on the continuous relaxations of the output variables to fin…