← Search

Cordelia Schmid

157 accepted papers

2026

CURVE: A Benchmark for Cultural and Multilingual Long Video Reasoning

CVPR 2026

Recent advancements in video models have shown tremendous progress, particularly in long video understanding. However, current benchmarks predominantly feature western-centric data and English as the dominant language, introducing significant biases in evaluation. To address this, we introduce CURVE

Cited by 0SourceScholar
2026

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects

CVPR 2026

Dense Video Object Captioning (DVOC) is the task of jointly detecting, tracking, and captioning object trajectories in a video, requiring the ability to understand spatio-temporal details and describe them in natural language.Due to the complexity of the task and the high cost associated with manual

Cited by 1SourceScholar
2026

MetricNet: Recovering Metric Scale in Generative Navigation Policies

ICRA 2026poster

Generative navigation policies have made rapid progress in improving end-to-end learned navigation. Despite their promising results, this paradigm has two structural problems. First, the sampled trajectories exist in an abstract, unscaled space without metric grounding. Second, the control strategy …

2026

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding

CVPR 2026

Video reasoning models are a core component of egocentric and embodied agents. However, standard benchmarks for assessing models provide only evaluation of the output (e.g. the answer to a question), without evaluation of inter- mediate reasoning steps, and most provide answers only in the text doma

Cited by 0SourcecodeScholar
2026

PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction

RSS 2026poster

Vision–Language–Action (VLA) models have shown strong potential for general-purpose robotic manipulation by leveraging large pretrained vision-language backbones. However, most existing VLAs rely primarily on 2D visual representations, which limits their ability to reason about fine-grained geometry…

Cited by 0SourceScholar
2026

VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation

CVPR 2026

Training vision-language models (VLMs) for complex reasoning remains a challenging task, i.a. due to the scarcity of high-quality image-text reasoning data. Conversely, text-based reasoning resources are abundant and scalable, but it is still an open question how to leveraging them for VLM reasoning

Cited by 0SourcecodeScholar
2026

What Are You Doing? A Closer Look at Controllable Human Video Generation

CVPR 2026

High-quality benchmarks are crucial for driving progress in machine learning research. However, despite the growing interest in video generation, there is no comprehensive dataset to evaluate human synthesis. Humans can perform a wide variety of actions and interactions, but existing datasets, like

Cited by 0SourcecodeScholar
2025

Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs

CVPR 2025poster

We address the task of video chaptering, i.e., partitioning a long video timeline into semantic units and generating corresponding chapter titles. While relatively underexplored, automatic chaptering has the potential to enable efficient navigation and content retrieval in long-form videos. In this…

Cited by 0SourcePDFScholar
2025

Dense Video Object Captioning from Disjoint Supervision

ICLR 2025spotlight

We propose a new task and model for dense video object captioning -- detecting, tracking and captioning trajectories of objects in a video. This task unifies spatial and temporal localization in video, whilst also requiring fine-grained visual understanding that is best described by natural language…

2025

FirePlace: Geometric Refinements of LLM Common Sense Reasoning for 3D Object Placement

CVPR 2025highlight

Scene generation with 3D assets presents a complex challenge, requiring both high-level semantic understanding and low-level geometric reasoning. While Multimodal Large Language Models (MLLMs) excel at semantic tasks, their application to 3D scene generation is hindered by their limited grounding on…

Cited by 2SourcePDFScholar
2025

Flexible Frame Selection for Efficient Video Reasoning

CVPR 2025poster

Video-language models have shown promise for addressing a range of multimodal tasks for video understanding, such as video question-answering. However, the inherent computational challenges of processing long video data and increasing model sizes have led to standard approaches that are limited by t…

Cited by 0SourcePDFScholar
2025

FlowNav: Combining Flow Matching and Depth Priors for Efficient Navigation

IROS 2025

Effective robot navigation in unseen environments is a challenging task that requires precise control actions at high frequencies. Recent advances have framed it as an image-goal-conditioned control problem, where the robot generates navigation actions using frontal RGB images. Current state-of-the-

Cited by 8SourceScholar
2025

HORT: Monocular Hand-held Objects Reconstruction with Transformers

ICCV 2025poster

Reconstructing hand-held objects in 3D from monocular images remains a significant challenge in computer vision. Most existing approaches rely on implicit 3D representations, which produce overly smooth reconstructions and are time-consuming to generate explicit 3D shapes. While more recent methods…

Cited by 0SourcePDFScholar
2025

InteractVLM: 3D Interaction Reasoning from 2D Foundational Models

CVPR 2025poster

We introduce InteractVLM, a novel method to estimate 3D contact points on human bodies and objects from single in-the-wild images, enabling accurate human-object joint reconstruction in 3D. This is challenging due to occlusions, depth ambiguities, and widely varying object shapes. Existing methods r…

2025

Language-Guided Image Tokenization for Generation

CVPR 2025poster

Image tokenization, the process of transforming raw image pixels into a compact low-dimensional latent representation, has proven crucial for scalable and efficient image generation. However, mainstream image tokenization methods generally have limited compression rates, making high-resolution image…

Cited by 7SourcePDFScholar
2025

Large-scale Pre-training for Grounded Video Caption Generation

ICCV 2025poster

We propose a novel approach for captioning and object grounding in video, where the objects in the caption are grounded in the video via temporally dense bounding boxes. We introduce the following contributions. First, we present a large-scale automatic annotation method that aggregates frame-level…

2025

MINERVA: Evaluating Complex Video Reasoning

ICCV 2025poster

Multimodal LLMs are turning their focus to video benchmarks, however most video benchmarks only provide outcome supervision, with no intermediate or interpretable reasoning steps. This makes it challenging to assess if models are truly able to combine perceptual and temporal information to reason ab…

2025

OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models

EMNLP 2025

Large vision-language models (VLMs) often struggle to generate long and factual captions. However, traditional measures for hallucination and factuality are not well suited for evaluating longer, more diverse captions and in settings where ground-truth human-annotated captions are unavailable. We in

Cited by 0SourcePDFScholar
2025

Prediction-Powered Causal Inferences

NeurIPS 2025poster

In many scientific experiments, the data annotating cost constraints the pace for testing novel hypotheses. Yet, modern machine learning pipelines offer a promising solution—provided their predictions yield correct conclusions. We focus on Prediction-Powered Causal Inferences (PPCI), i.e., estimatin…

Cited by 0SourceScholar
2025

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

NeurIPS 2025poster

Despite recent advances in Vision-Language Models (VLMs), long-video understanding remains a challenging problem. Although state-of-the-art long-context VLMs can process around 1000 input frames, they still struggle to effectively leverage this sequence length, and succumb to irrelevant distractors…

Cited by 0SourceScholar
2025

Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-Guided 3D Policy

ICRA 2025

Generalizing language-conditioned robotic policies to new tasks remains a significant challenge, hampered by the lack of suitable simulation benchmarks. In this paper, we address this gap by introducing GemBench, a novel benchmark to assess generalization capabilities of vision-language robotic mani

Cited by 40SourcecodeScholar
2025

Towards Zero-Shot Multimodal Machine Translation

NAACL 2025findings

Current multimodal machine translation (MMT) systems rely on fully supervised data (i.e sentences with their translations and accompanying images), which is costly to collect and prevents the extension of MMT to language pairs with no such data. We propose a method to bypass the need for fully super…

2025

ViViDex: Learning Vision-Based Dexterous Manipulation from Human Videos

ICRA 2025

In this work, we aim to learn a unified vision-based policy for multi-fingered robot hands to manipulate a variety of objects in diverse poses. Though prior work has shown benefits of using human videos for policy learning, performance gains have been limited by the noise in estimated trajectories.

Cited by 36SourcecodeScholar
2025

Visual Lexicon: Rich Image Features in Language Space

CVPR 2025poster

We present Visual Lexicon, a novel visual language that encodes rich image information into the text space of vocabulary tokens while retaining intricate visual details that are often challenging to convey in natural language. Unlike traditional methods that prioritize either high-level semantics (e…

Cited by 1SourcePDFScholar
2025

mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus

ACL 2025finding

Multimodal Large Language Models (mLLMs) are trained on a large amount of text-image data. While most mLLMs are trained on caption-like data only, Alayrac et al. (2022) showed that additionally training them on interleaved sequences of text and images can lead to the emergence of in-context learning…

2024

A Generative Approach for Wikipedia-Scale Visual Entity Recognition

CVPR 2024poster

In this paper we address web-scale visual entity recognition specifically the task of mapping a given query image to one of the 6 million existing entities in Wikipedia. One way of approaching a problem of such scale is using dual encoder models (e.g. CLIP) where all the entity names and query image…

2024

CoVR: Learning Composed Video Retrieval from Web Video Captions

AAAI 2024technical

Composed Image Retrieval (CoIR) has recently gained popularity as a task that considers both text and image queries together, to search for relevant images in a database. Most CoIR approaches require manually annotated datasets, comprising image-text-image triplets, where the text describes a modifi…

Cited by 46SourcePDFScholar
2024

DataDream: Few-shot Guided Dataset Generation

ECCV 2024poster

"While text-to-image diffusion models have been shown to achieve state-of-the-art results in image synthesis, they have yet to prove their effectiveness in downstream applications. Previous work has proposed to generate data for image classifier training given limited real data access. However, thes…

2024

End-to-End Spatio-Temporal Action Localisation with Video Transformers

CVPR 2024poster

The most performant spatio-temporal action localisation models use external person proposals and complex external memory banks. We propose a fully end-to-end transformer based model that directly ingests an input video and outputs tubelets -- a sequence of bounding boxes and the action classes at ea…

Cited by 21SourcePDFScholar
2024

Learning Correlation Structures for Vision Transformers

CVPR 2024poster

We introduce a new attention mechanism dubbed structural self-attention (StructSA) that leverages rich correlation patterns naturally emerging in key-query interactions of attention. StructSA generates attention maps by recognizing space-time structures of key-query correlations via convolution and…

Cited by 19SourcePDFScholar
2024

MoReVQA: Exploring Modular Reasoning Models for Video Question Answering

CVPR 2024poster

This paper addresses the task of video question answering (videoQA) via a decomposed multi-stage modular reasoning framework. Previous modular methods have shown promise with a single planning stage ungrounded in visual content. However through a simple and effective baseline we find that such syste…

Cited by 32SourcePDFScholar
2024

SUGAR: Pre-training 3D Visual Representations for Robotics

CVPR 2024poster

Learning generalizable visual representations from Internet data has yielded promising results for robotics. Yet prevailing approaches focus on pre-training 2D representations being sub-optimal to deal with occlusions and accurately localize objects in complex 3D scenes. Meanwhile 3D representation…

Cited by 15SourcePDFScholar
2024

SceneCraft: An LLM Agent for Synthesizing 3D Scenes as Blender Code

ICML 2024oral

This paper introduces SceneCraft, a Large Language Model (LLM) Agent converting text descriptions into Blender-executable Python scripts which render complex scenes with up to a hundred 3D assets. This process requires complex spatial planning and arrangement. We tackle these challenges through a co…

Cited by 39SourcePDFScholar
2024

Smoke and Mirrors in Causal Downstream Tasks

NeurIPS 2024poster

Machine Learning and AI have the potential to transform data-driven scientific discovery, enabling accurate predictions for several scientific phenomena. As many scientific questions are inherently causal, this paper looks at the causal inference task of treatment effect estimation, where the outcom…

2024

Streaming Dense Video Captioning

CVPR 2024poster

An ideal model for dense video captioning -- predicting captions localized temporally in a video -- should be able to handle long input videos predict rich detailed textual descriptions and be able to produce outputs before processing the entire video. Current state-of-the-art models however process…

2024

Time- Memory- and Parameter-Efficient Visual Adaptation

CVPR 2024highlight

As foundation models become more popular there is a growing need to efficiently finetune them for downstream tasks. Although numerous adaptation methods have been proposed they are designed to be efficient only in terms of how many parameters are trained. They however typically still require backpro…

Cited by 15SourcePDFScholar
2024

Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach

NeurIPS 2024poster

Web-scale visual entity recognition, the task of associating images with their corresponding entities within vast knowledge bases like Wikipedia, presents significant challenges due to the lack of clean, large-scale training data. In this paper, we propose a novel methodology to curate such a datase…

Cited by 2SourcePDFScholar
2023

AVFormer: Injecting Vision Into Frozen Speech Models for Zero-Shot AV-ASR

CVPR 2023poster

Audiovisual automatic speech recognition (AV-ASR) aims to improve the robustness of a speech recognition system by incorporating visual information. Training fully supervised multimodal models for this task from scratch, however is limited by the need for large labelled audiovisual datasets (in each…

Cited by 15SourcePDFScholar
2023

AVIS: Autonomous Visual Information Seeking with Large Language Model Agent

NeurIPS 2023poster

In this paper, we propose an autonomous information seeking visual question answering framework, AVIS. Our method leverages a Large Language Model (LLM) to dynamically strategize the utilization of external tools and to investigate their outputs via tree search, thereby acquiring the indispensable k…

Cited by 51SourcePDFScholar
2023

Audiovisual Masked Autoencoders

ICCV 2023poster

Can we leverage the audiovisual information already present in video to improve self-supervised representation learning? To answer this question, we study various pretraining architectures and objectives within the masked autoencoding framework, motivated by the success of similar methods in natural…

Cited by 57PDFcodeScholar
2023

Bridging the Gap Between Model Explanations in Partially Annotated Multi-Label Classification

CVPR 2023poster

Due to the expensive costs of collecting labels in multi-label classification datasets, partially annotated multi-label classification has become an emerging field in computer vision. One baseline approach to this task is to assume unobserved labels as negative labels, but this assumption induces la…

2023

Does Visual Pretraining Help End-to-End Reasoning?

NeurIPS 2023poster

We aim to investigate whether end-to-end learning of visual reasoning can be achieved with general-purpose neural networks, with the help of visual pretraining. A positive result would refute the common belief that explicit visual abstraction (e.g. object detection) is essential for compositional ge…

Cited by 4SourcePDFScholar
2023

Enforcing the consensus between Trajectory Optimization and Policy Learning for precise robot control

ICRA 2023poster

Reinforcement learning (RL) and trajectory opti-mization (TO) present strong complementary advantages. On one hand, RL approaches are able to learn global control policies directly from data, but generally require large sample sizes to properly converge towards feasible policies. On the other hand,…

Cited by 7SourceScholar
2023

Improving Image Recognition by Retrieving From Web-Scale Image-Text Data

CVPR 2023poster

Retrieval augmented models are becoming increasingly popular for computer vision tasks after their recent success in NLP problems. The goal is to enhance the recognition capabilities of the model by retrieving similar examples for the visual input from an external memory set. In this work, we introd…

Cited by 27SourcePDFScholar
2023

Learning Reward Functions for Robotic Manipulation by Observing Humans

ICRA 2023poster

Observing a human demonstrator manipulate objects provides a rich, scalable and inexpensive source of data for learning robotic policies. However, transferring skills from human videos to a robotic manipulator poses several challenges, not least a difference in action and observation spaces. In this…

Cited by 26SourceScholar
2023

Modular Visual Question Answering via Code Generation

ACL 2023short

We present a framework that formulates visual question answering as modular code generation. In contrast to prior work on modular approaches to VQA, our approach requires no additional training and relies on pre-trained language models (LMs), visual models pre-trained on image-caption pairs, and fif…

2023

PolarNet: 3D Point Clouds for Language-Guided Robotic Manipulation

CoRL 2023poster

The ability for robots to comprehend and execute manipulation tasks based on natural language instructions is a long-term goal in robotics. The dominant approaches for language-guided manipulation use 2D image representations, which face difficulties in combining multi-view cameras and inferring pre…

Cited by 37SourcecodeScholar
2023

REVEAL: Retrieval-Augmented Visual-Language Pre-Training With Multi-Source Multimodal Knowledge Memory

CVPR 2023highlight

In this paper, we propose an end-to-end Retrieval-Augmented Visual Language Model (REVEAL) that learns to encode world knowledge into a large-scale memory, and to retrieve from it to answer knowledge-intensive queries. REVEAL consists of four key components: the memory, the encoder, the retriever an…

2023

Robust Visual Sim-to-Real Transfer for Robotic Manipulation

IROS 2023poster

Learning visuomotor policies in simulation is much safer and cheaper than in the real world. However, due to discrepancies between the simulated and real data, simulator-trained policies often fail when transferred to real robots. One common approach to bridge the visual sim-to-real domain gap is do…

Cited by 4SourceScholar
2023

Tackling Ambiguity with Images: Improved Multimodal Machine Translation and Contrastive Evaluation

ACL 2023long

One of the major challenges of machine translation (MT) is ambiguity, which can in some cases be resolved by accompanying context such as images. However, recent work in multimodal MT (MMT) has shown that obtaining improvements from images is challenging, limited not only by the difficulty of buildi…

2023

UnLoc: A Unified Framework for Video Localization Tasks

ICCV 2023poster

While large-scale image-text pretrained models such as CLIP have been used for multiple video-level tasks on trimmed videos, their use for temporal localization in untrimmed videos is still a relatively unexplored task. We design a new approach for this called UnLoc, which uses pretrained image and…

Cited by 61PDFcodeScholar
2023

Verbs in Action: Improving Verb Understanding in Video-Language Models

ICCV 2023poster

Understanding verbs is crucial to modelling how people and objects interact with each other and the environment through space and time. Recently, state-of-the-art video-language models based on CLIP have been shown to have limited verb understanding and to rely extensively on nouns, restricting thei…

Cited by 82PDFcodeScholar
2023

Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning

CVPR 2023poster

In this work, we introduce Vid2Seq, a multi-modal single-stage dense event captioning model pretrained on narrated videos which are readily-available at scale. The Vid2Seq architecture augments a language model with special time tokens, allowing it to seamlessly predict event boundaries and textual…

2023

VidChapters-7M: Video Chapters at Scale

NeurIPS 2023poster

Segmenting untrimmed videos into chapters enables users to quickly navigate to the information of their interest. This important topic has been understudied due to the lack of publicly released datasets. To address this issue, we present VidChapters-7M, a dataset of 817K user-chaptered videos includ…

Cited by 36SourcePDFScholar
2023

WALDO: Future Video Synthesis Using Object Layer Decomposition and Parametric Flow Prediction

ICCV 2023poster

This paper presents WALDO (WArping Layer-Decomposed Objects), a novel approach to the prediction of future video frames from past ones. Individual images are decomposed into multiple layers combining object masks and a small set of control points. The layer structure is shared across all frames in e…

Cited by 8PDFcodeScholar
2023

Waffling Around for Performance: Visual Classification with Random Words and Broad Concepts

ICCV 2023poster

The visual classification performance of vision-language models such as CLIP has been shown to benefit from additional semantic knowledge from large language models (LLMs) such as GPT-3. In particular, averaging over LLM-generated class descriptors, e.g. "waffle, which has a round shape", can notabl…

Cited by 86PDFcodeScholar
2023

gSDF: Geometry-Driven Signed Distance Functions for 3D Hand-Object Reconstruction

CVPR 2023poster

Signed distance functions (SDFs) is an attractive framework that has recently shown promising results for 3D shape reconstruction from images. SDFs seamlessly generalize to different shape resolutions and topologies but lack explicit modelling of the underlying 3D geometry. In this work, we exploit…

2022

AlignSDF: Pose-Aligned Signed Distance Fields for Hand-Object Reconstruction

ECCV 2022poster

"Recent work achieved impressive progress towards joint reconstruction of hands and manipulated objects from monocular color images. Existing methods focus on two alternative representations in terms of either parametric meshes or signed distance fields (SDFs). On one side, parametric models can ben…

2022

Assembly Planning from Observations under Physical Constraints

IROS 2022poster

This paper addresses the problem of copying an unknown assembly of primitives with known shape and appearance using information extracted from a single photograph by an off-the-shelf procedure for object detection and pose estimation. The proposed algorithm uses a simple combination of physical stab…

Cited by 5SourceScholar
2022

End-to-End Generative Pretraining for Multimodal Video Captioning

CVPR 2022poster

Recent video and language pretraining frameworks lack the ability to generate sentences. We present Multimodal Video Generative Pretraining (MV-GPT), a new pretraining framework for learning from unlabelled videos which can be effectively used for generative tasks such as multimodal video captioning…

Cited by 220PDFScholar
2022

Instruction-driven history-aware policies for robotic manipulations

CoRL 2022oral

In human environments, robots are expected to accomplish a variety of manipulation tasks given simple natural language instructions. Yet, robotic manipulation is extremely challenging as it requires fine-grained motor control, long-term memory as well as generalization to previously unseen tasks and…

Cited by 117SourcecodeScholar
2022

Language Conditioned Spatial Relation Reasoning for 3D Object Grounding

NeurIPS 2022accept

Localizing objects in 3D scenes based on natural language requires understanding and reasoning about spatial relations. In particular, it is often crucial to distinguish similar objects referred by the text, such as "the left most chair" and "a chair next to the window". In this work we propose a la…

2022

Learning Audio-Video Modalities from Image Captions

ECCV 2022poster

"There has been a recent explosion of large-scale image-text datasets, as images with alt-text captions can be easily obtained online. Obtaining large-scale, high quality data for video in the form of text-video and text-audio pairs however, is more challenging. To close this gap we propose a new vi…

Cited by 109SourcePDFScholar
2022

Learning from Unlabeled 3D Environments for Vision-and-Language Navigation

ECCV 2022poster

"In vision-and-language navigation (VLN), an embodied agent is required to navigate in realistic 3D environments following natural language instructions. One major bottleneck for existing VLN approaches is the lack of sufficient training data, resulting in unsatisfactory generalization to unseen env…

2022

Multiview Transformers for Video Recognition

CVPR 2022poster

Video understanding requires reasoning at multiple spatiotemporal resolutions -- from short fine-grained motions to events taking place over longer durations. Although transformer architectures have recently advanced the state-of-the-art, they have not explicitly modelled different spatiotemporal re…

Cited by 348PDFcodeScholar
2022

TL;DW? Summarizing Instructional Videos with Task Relevance & Cross-Modal Saliency

ECCV 2022poster

"YouTube users looking for instructions for a specific task may spend a long time browsing content trying to find the right video that matches their needs. Creating a visual summary (abridged version of a video) provides viewers with a quick overview and massively reduces search time. In this work,…

Cited by 0SourcePDFScholar
2022

Think Global, Act Local: Dual-Scale Graph Transformer for Vision-and-Language Navigation

CVPR 2022oral

Following language instructions to navigate in unseen environments is a challenging problem for autonomous embodied agents. The agent not only needs to ground languages in visual scenes, but also should explore the environment to reach its target. In this work, we propose a dual-scale graph transfor…

Cited by 181PDFScholar
2022

TubeDETR: Spatio-Temporal Video Grounding With Transformers

CVPR 2022oral

We consider the problem of localizing a spatio-temporal tube in a video corresponding to a given text query. This is a challenging task that requires the joint and efficient modeling of temporal, spatial and multi-modal interactions. To address this task, we propose TubeDETR, a transformer-based arc…

Cited by 120PDFcodeScholar
2022

Zero-Shot Video Question Answering via Frozen Bidirectional Language Models

NeurIPS 2022accept

Video question answering (VideoQA) is a complex task that requires diverse multi-modal data for training. Manual annotation of question and answers for videos, however, is tedious and prohibits scalability. To tackle this problem, recent methods consider zero-shot settings with no manual annotation…

2021

Airbert: In-Domain Pretraining for Vision-and-Language Navigation

ICCV 2021poster

Vision-and-language navigation (VLN) aims to enable embodied agents to navigate in realistic environments using natural language instructions. Given the scarcity of domain-specific training data and the high diversity of image and language inputs, the generalization of VLN agents to unseen environme…

Cited by 169PDFcodeScholar
2021

Attention Bottlenecks for Multimodal Fusion

NeurIPS 2021poster

Humans perceive the world by concurrently processing and fusing high-dimensional inputs from multiple modalities such as vision and audio. Machine perception models, in stark contrast, are typically modality-specific and optimised for unimodal benchmarks. A common approach for building multimodal m…

2021

Composable Augmentation Encoding for Video Representation Learning

ICCV 2021poster

We focus on contrastive methods for self-supervised video representation learning. A common paradigm in contrastive learning is to construct positive pairs by sampling different data views for the same instance, with different data instances as negatives. These methods implicitly assume a set of rep…

Cited by 26PDFcodeScholar
2021

Differentiable Simulation for Physical System Identification

RA-L 2021

Simulating frictional contacts remains a challenging research topic in robotics. Recently, differentiable physics emerged and has proven to be a key element in model-based Reinforcement Learning (RL) and optimal control fields. However, most of the current formulations deploy coarse approximations o

Cited by 65SourceScholar
2021

Differentiable rendering with perturbed optimizers

NeurIPS 2021poster

Reasoning about 3D scenes from their 2D image projections is one of the core problems in computer vision. Solutions to this inverse and ill-posed problem typically involve a search for models that best explain observed image data. Notably, images depend both on the properties of observed scenes and…

Cited by 16SourcePDFScholar
2021

Goal-Conditioned Reinforcement Learning with Imagined Subgoals

ICML 2021spotlight

Goal-conditioned reinforcement learning endows an agent with a large variety of skills, but it often struggles to solve tasks that require more temporally extended reasoning. In this work, we propose to incorporate imagined subgoals into policy learning to facilitate learning of complex tasks. Imagi…

Cited by 170SourcePDFScholar
2021

HDMapGen: A Hierarchical Graph Generative Model of High Definition Maps

CVPR 2021poster

High Definition (HD) maps are maps with precise definitions of road lanes with rich semantics of the traffic rules. They are critical for several key stages in an autonomous driving system, including motion forecasting and planning. However, there are only a small amount of real-world road topologie…

Cited by 71PDFScholar
2021

History Aware Multimodal Transformer for Vision-and-Language Navigation

NeurIPS 2021poster

Vision-and-language navigation (VLN) aims to build autonomous visual agents that follow instructions and navigate in real scenes. To remember previously visited locations and actions taken, most approaches to VLN implement memory using recurrent states. Instead, we introduce a History Aware Multimod…

Cited by 268SourcePDFScholar
2021

Just Ask: Learning To Answer Questions From Millions of Narrated Videos

ICCV 2021poster

Recent methods for visual question answering rely on large-scale annotated datasets. Manual annotation of questions and answers for videos, however, is tedious, expensive and prevents scalability. In this work, we propose to avoid manual annotation and generate a large-scale training dataset for vid…

Cited by 352PDFcodeScholar
2021

Large-Scale Unsupervised Object Discovery

NeurIPS 2021poster

Existing approaches to unsupervised object discovery (UOD) do not scale up to large datasets without approximations that compromise their performance. We propose a novel formulation of UOD as a ranking problem, amenable to the arsenal of distributed methods available for eigenvalue problems and link…

2021

ViViT: A Video Vision Transformer

ICCV 2021poster

We present pure-transformer based models for video classification, drawing upon the recent success of such models in image classification. Our model extracts spatio-temporal tokens from the input video, which are then encoded by a series of transformer layers. In order to handle the long sequences o…

Cited by 2888PDFcodeScholar
2020

Ava Active Speaker: An Audio-Visual Dataset for Active Speaker Detection

ICASSP 2020accepted

Active speaker detection is an important component in video analysis algorithms for applications such as speaker diarization, video re-targeting for meetings, speech enhancement, and human-robot interaction. The absence of a large, carefully labeled audio-visual active speaker dataset has limited ev…

Cited by 0SourceScholar
2020

Graph convolutional networks for learning with few clean and many noisy labels

ECCV 2020poster

In this work we consider the problem of learning a classifier from noisy labels when a few clean labeled examples are given. The structure of clean and noisy data is modeled by a graph per class and Graph Convolutional Networks (GCN) are used to predict class relevance of noisy examples. For each cl…

2020

Learning Obstacle Representations for Neural Motion Planning

CoRL 2020

Motion planning and obstacle avoidance is a key challenge in robotics applications. While previous work succeeds to provide excellent solutions for known environments, sensor-based motion planning in new and dynamic environments remains to be difficult. In this work we address sensor-based motion pl

2020

Learning to combine primitive skills: A step towards versatile robotic manipulation §

ICRA 2020poster

Manipulation tasks such as preparing a meal or assembling furniture remain highly challenging for robotics and vision. Traditional task and motion planning (TAMP) methods can solve complex tasks but require full state observability and are not adapted to dynamic scene changes. Recent learning method…

Cited by 56SourcecodeScholar
2020

Learning visual policies for building 3D shape categories

IROS 2020poster

Manipulation and assembly tasks require non-trivial planning of actions depending on the environment and the final goal. Previous work in this domain often assembles particular instances of objects from known sets of primitives. In contrast, we aim to handle varying sets of primitives and to constru…

Cited by 6SourceScholar
2020

Leveraging Photometric Consistency Over Time for Sparsely Supervised Hand-Object Reconstruction

CVPR 2020poster

Modeling hand-object manipulations is essential for understanding how humans interact with their environment. While of practical importance, estimating the pose of hands and objects during interactions is challenging due to the large mutual occlusions that occur during manipulation. Recent efforts h…

Cited by 217PDFScholar
2020

Memory-Efficient Incremental Learning Through Feature Adaptation

ECCV 2020poster

We introduce an approach for incremental learning that preserves feature descriptors of training images from previously learned classes, instead of the images themselves, unlike most existing work. Keeping the much lower-dimensional feature embeddings of images reduces the memory footprint significa…

Cited by 226SourcePDFScholar
2020

Radioactive data: tracing through training

ICML 2020poster

Data tracing determines whether particular data samples have been used to train a model. We propose a new technique, radioactive data, that makes imperceptible changes to these samples such that any model trained on them will bear an identifiable mark. Given a trained model, our technique detects th…

2020

Selecting Relevant Features from a Multi-domain Representation for Few-shot Classification

ECCV 2020poster

Popular approaches for few-shot classification consist of first learning a generic data representation based on a large annotated dataset, before adapting the representation to new classes given only a few labeled samples. In this work, we propose a new strategy based on feature selection, which is…

2020

Speech2Action: Cross-Modal Supervision for Action Recognition

CVPR 2020poster

Is it possible to guess human action from dialogue alone? In this work we investigate the link between spoken words and actions in movies. We note that movie screenplays describe actions, as well as contain the speech of characters and hence can be used to learn this correlation with no additional s…

Cited by 78PDFScholar
2020

TAO: A Large-Scale Benchmark for Tracking Any Object

ECCV 2020poster

For many years, multi-object tracking benchmarks have focused on a handful of categories. Motivated primarily by surveillance and self-driving applications, these datasets provide tracks for people, vehicles, and animals, ignoring the vast majority of objects in the world. By contrast, in the relate…

Cited by 214SourcePDFScholar
2020

Uncertainty-Aware Weakly Supervised Action Detection from Untrimmed Videos

ECCV 2020poster

Despite the recent advances in video classification, progress in spatio-temporal action recognition has lagged behind. A major contributing factor has been the prohibitive cost of annotating videos frame-by-frame. In this paper, we present a spatio-temporal action recognition model that is trained w…

2020

VectorNet: Encoding HD Maps and Agent Dynamics From Vectorized Representation

CVPR 2020poster

Behavior prediction in dynamic, multi-agent systems is an important problem in the context of self-driving cars, due to the complex representations and interactions of road components, including moving agents (e.g. pedestrians and vehicles) and road context information (e.g. lanes, traffic lights).…

Cited by 1022PDFScholar
2020

What Makes for Good Views for Contrastive Learning?

NeurIPS 2020poster

Contrastive learning between multiple views of the data has recently achieved state of the art performance in the field of self-supervised representation learning. Despite its success, the influence of different view choices has been less studied. In this paper, we use theoretical and empirical anal…

Cited by 1641SourcePDFScholar
2019

Adaptive Density Estimation for Generative Models

NeurIPS 2019spotlight

Unsupervised learning of generative models has seen tremendous progress over recent years, in particular due to generative adversarial networks (GANs), variational autoencoders, and flow-based models. GANs have dramatically improved sample quality, but suffer from two drawbacks: (i) they mode-drop,…

Cited by 30SourcePDFScholar
2019

Learning Joint Reconstruction of Hands and Manipulated Objects

CVPR 2019poster

Estimating hand-object manipulations is essential for in- terpreting and imitating human actions. Previous work has made significant progress towards reconstruction of hand poses and object shapes in isolation. Yet, reconstructing hands and objects during manipulation is a more challeng- ing task du…

Cited by 631PDFScholar
2019

Learning to Augment Synthetic Images for Sim2Real Policy Transfer

IROS 2019poster

Vision and learning have made significant progress that could improve robotics policies for complex tasks and environments. Learning deep neural networks for image understanding, however, requires large amounts of domain-specific visual data. While collecting such data from real robots is possible,…

Cited by 55SourcecodeScholar
2019

MARS: Motion-Augmented RGB Stream for Action Recognition

CVPR 2019poster

Most state-of-the-art methods for action recognition consist of a two-stream architecture with 3D convolutions: an appearance stream for RGB frames and a motion stream for optical flow frames. Although combining flow with RGB improves the performance, the cost of computing accurate optical flow is…

Cited by 336PDFScholar
2019

Moulding Humans: Non-Parametric 3D Human Shape Estimation From Single Images

ICCV 2019poster

In this paper, we tackle the problem of 3D human shape estimation from single RGB images. While the recent progress in convolutional neural networks has allowed impressive results for 3D human pose estimation, estimating the full 3D shape of a person is still an open issue. Model-based approaches ca…

Cited by 152PDFScholar
2019

Relational Action Forecasting

CVPR 2019oral

This paper focuses on multi-person action forecasting in videos. More precisely, given a history of H previous frames, the goal is to detect actors and to predict their future actions for the next T frames. Our approach jointly models temporal and spatial interactions among different actors by const…

Cited by 100PDFScholar
2019

Self-Supervised Learning With Geometric Constraints in Monocular Video: Connecting Flow, Depth, and Camera

ICCV 2019poster

We present GLNet, a self-supervised framework for learning depth, optical flow, camera pose and intrinsic parameters from monocular video -- addressing the difficulty of acquiring realistic ground-truth for such tasks. We propose three contributions: 1) we design new loss functions that capture mult…

Cited by 317PDFScholar
2019

Spreading vectors for similarity search

ICLR 2019poster

Discretizing floating-point vectors is a fundamental step of modern indexing methods. State-of-the-art techniques learn parameters of the quantizers on training data for optimal performance, thus adapting quantizers to the data. In this work, we propose to reverse this paradigm and adapt the data to…

2019

VideoBERT: A Joint Model for Video and Language Representation Learning

ICCV 2019poster

Self-supervised learning has become increasingly important to leverage the abundance of unlabeled data available on platforms like YouTube. Whereas most existing approaches learn low-level representations, we propose a joint visual-linguistic model to learn high-level features without any explicit s…

Cited by 1568PDFScholar
2019

White-box vs Black-box: Bayes Optimal Strategies for Membership Inference

ICML 2019oral

Membership inference determines, given a sample and trained parameters of a machine learning model, whether the sample was part of the training set. In this paper, we derive the optimal strategy for membership inference with a few assumptions on the distribution of the parameters. We show that optim…

Cited by 432SourcePDFScholar
2018

A flexible model for training action localization with varying levels of supervision

NeurIPS 2018poster

Spatio-temporal action detection in videos is typically addressed in a fully-supervised setup with manual annotation of training videos required at every frame. Since such annotation is extremely tedious and prohibits scalability, there is a clear need to minimize the amount of manual supervision.…

2018

AVA: A Video Dataset of Spatio-Temporally Localized Atomic Visual Actions

CVPR 2018poster

This paper introduces a video dataset of spatio-temporally localized Atomic Visual Actions (AVA). The AVA dataset densely annotates 80 atomic visual actions in 437 15-minute video clips, where actions are localized in space and time, resulting in 1.59M action labels with multiple labels per person o…

Cited by 1319SourcePDFScholar
2018

Actor and Observer: Joint Modeling of First and Third-Person Videos

CVPR 2018poster

Several theories in cognitive neuroscience suggest that when people interact with the world, or simulate interactions, they do so from a first-person egocentric perspective, and seamlessly transfer knowledge between third-person (observer) and first-person (actor). Despite this, learning such models…

2018

Actor-centric Relation Network

ECCV 2018poster

Current state-of-the-art approaches for spatio-temporal action localization rely on detections at the frame level and model temporal context with 3D ConvNets. Here, we go one step further and model spatio-temporal relations to capture the interactions between human actors, relevant objects and scene…

Cited by 280SourcePDFScholar
2018

BodyNet: Volumetric Inference of 3D Human Body Shapes

ECCV 2018poster

Human shape estimation is an important task for video editing, animation and fashion industry. Predicting 3D human body shape from natural images, however, is highly challenging due to factors such as variation in human bodies, clothing and viewpoint. Prior methods addressing this problem typically…

Cited by 530SourcePDFScholar
2018

End-to-End Incremental Learning

ECCV 2018poster

Although deep learning approaches have stood out in recent years due to their state-of-the-art results, they continue to suffer from catastrophic forgetting, a dramatic decrease in overall performance when training with new classes added incrementally. This is due to current neural network architect…

2018

Modeling Visual Context is Key to Augmenting Object Detection Datasets

ECCV 2018poster

Performing data augmentation for learning deep neural networks is well known to be important for training visual recognition systems. By artificially increasing the number of training examples, it helps reducing overfitting and improves generalization. For object detection, classical approaches for…

Cited by 314SourcePDFScholar
2018

PoTion: Pose MoTion Representation for Action Recognition

CVPR 2018poster

Most state-of-the-art methods for action recognition rely on a two-stream architecture that processes appearance and motion independently. In this paper, we claim that considering them jointly offers rich information for action recognition. We introduce a novel representation that gracefully encodes…

Cited by 381SourcePDFScholar
2018

Unsupervised Learning of Artistic Styles with Archetypal Style Analysis

NeurIPS 2018poster

In this paper, we introduce an unsupervised learning approach to automatically dis- cover, summarize, and manipulate artistic styles from large collections of paintings. Our method is based on archetypal analysis, which is an unsupervised learning technique akin to sparse coding with a geometric int…

2017

Action Tubelet Detector for Spatio-Temporal Action Localization

ICCV 2017poster

Current state-of-the-art approaches for spatio-temporal action localization rely on detections at the frame level that are then linked or tracked across time. In this paper, we leverage the temporal continuity of videos instead of operating at the frame level. We propose the ACtion Tubelet detector…

Cited by 437PDFcodeScholar
2017

BlitzNet: A Real-Time Deep Network for Scene Understanding

ICCV 2017poster

Real-time scene understanding has become crucial in many applications such as autonomous driving. In this paper, we propose a deep architecture, called BlitzNet, that jointly performs object detection and semantic segmentation in one forward pass, allowing real-time computations. Besides the computa…

Cited by 264PDFScholar
2017

Incremental Learning of Object Detectors Without Catastrophic Forgetting

ICCV 2017poster

Despite their success for object detection, convolutional neural networks are ill-equipped for incremental learning, i.e., adapting the original model trained on a set of classes to additionally detect objects of new classes, in the absence of the initial training data. They suffer from "catastrophi…

Cited by 685PDFScholar
2017

Learning From Synthetic Humans

CVPR 2017poster

Estimating human pose, shape, and motion from images and video are fundamental challenges with many applications. Recent advances in 2D human pose estimation use large amounts of manually-labeled training data for learning convolutional neural networks (CNNs). Such data is time consuming to acquire…

Cited by 1234PDFScholar
2017

SCNet: Learning Semantic Correspondence

ICCV 2017poster

This paper addresses the problem of establishing semantic correspondences between images depicting different instances of the same object or scene category. Previous approaches focus on either combining a spatial regularizer with hand-crafted features, or learning a correspondence model for appearan…

Cited by 159PDFcodeScholar
2016

Proposal Flow

CVPR 2016poster

Finding image correspondences remains a challenging problem in the presence of intra-class variations and large changes in scene layout. Semantic flow methods are designed to handle images depicting different instances of the same object or scene category. We introduce a novel approach to semantic…

Cited by 164PDFScholar
2015

EpicFlow: Edge-Preserving Interpolation of Correspondences for Optical Flow

CVPR 2015poster

We propose a novel approach for optical flow estimation, targeted at large displacements with significant occlusions. It consists of two steps: i) dense matching by edge-preserving interpolation from a sparse set of matches; ii) variational energy minimization initialized with the dense matches. The…

2015

Learning to Detect Motion Boundaries

CVPR 2015poster

We propose a learning-based approach for motion boundary detection. Precise localization of motion boundaries is essential for the success of optical flow estimation, as motion boundaries correspond to discontinuities of the optical flow field. The proposed approach allows to predict motion boundari…

2015

Local Convolutional Features With Unsupervised Training for Image Retrieval

ICCV 2015poster

Patch-level descriptors underlie several important computer vision tasks, such as stereo-matching or content-based image retrieval. We introduce a deep convolutional architecture that yields patch-level descriptors, as an alternative to the popular SIFT descriptor for image retrieval. The propo…

Cited by 219PDFScholar
2015

Unsupervised Object Discovery and Localization in the Wild: Part-Based Matching With Bottom-Up Region Proposals

CVPR 2015poster

This paper addresses unsupervised discovery and localization of dominant objects from a noisy image collection with multiple object classes. The setting of this problem is fully unsupervised, without even image-level annotations or any assumption of a single dominant class. This is far more general…

Cited by 321SourcePDFScholar
2015

Unsupervised Object Discovery and Tracking in Video Collections

ICCV 2015poster

This paper addresses the problem of automatically localizing dominant objects as spatio-temporal tubes in a noisy collection of videos with minimal or even no supervision. We formulate the problem as a combination of two complementary processes: discovery and tracking. The first one establishes corr…

Cited by 153PDFScholar
2015

Weakly-Supervised Alignment of Video With Text

ICCV 2015poster

Suppose that we are given a set of videos, along with natural language descriptions in the form of multiple sentences (e.g., manual annotations, movie scripts, sport summaries etc.), and that these sentences appear in the same temporal order as their visual counterparts. We propose in this paper a m…

Cited by 171PDFcodeScholar