← Search

Nikita Dvornik

13 accepted papers

2025

FOMO-3D: Using Vision Foundation Models for Long-Tailed 3D Object Detection

CoRL 2025poster

In order to navigate complex traffic environments, self-driving vehicles must recognize many semantic classes pertaining to vulnerable road users or traffic control devices. However, many safety-critical objects (e.g., construction worker) appear infrequently in nominal traffic conditions, leading…

Cited by 0SourceScholar
2023

GePSAn: Generative Procedure Step Anticipation in Cooking Videos

ICCV 2023poster

We study the problem of future step anticipation in procedural videos. Given a video of an ongoing procedural activity, we predict a plausible next procedure step described in rich natural language. While most previous work focus on the problem of data scarcity in procedural video datasets, another…

Cited by 11PDFcodeScholar
2023

LabelFormer: Object Trajectory Refinement for Offboard Perception from LiDAR Point Clouds

CoRL 2023poster

A major bottleneck to scaling-up training of self-driving perception systems are the human annotations required for supervision. A promising alternative is to leverage “auto-labelling” offboard perception models that are trained to automatically generate annotations from raw LiDAR point clouds at a…

Cited by 8SourceScholar
2023

Self-Supervised Learning of Action Affordances as Interaction Modes

ICRA 2023poster

When humans perform a task with an articulated object, they interact with the object only in a handful of ways, while the space of all possible interactions is nearly endless. This is because humans have prior knowledge about what interactions are likely to be successful, i.e., to open a new door we…

Cited by 5SourcecodeScholar
2023

SlotFormer: Unsupervised Visual Dynamics Simulation with Object-Centric Models

ICLR 2023poster

Understanding dynamics from visual observations is a challenging problem that requires disentangling individual objects from the scene and learning their interactions. While recent object-centric models can successfully decompose a scene into objects, modeling their dynamics effectively still remain…

2023

StepFormer: Self-Supervised Step Discovery and Localization in Instructional Videos

CVPR 2023poster

Instructional videos are an important resource to learn procedural tasks from human demonstrations. However, the instruction steps in such videos are typically short and sparse, with most of the video being irrelevant to the procedure. This motivates the need to temporally localize the instruction s…

Cited by 29SourcePDFScholar
2022

Flow Graph to Video Grounding for Weakly-Supervised Multi-step Localization

ECCV 2022poster

"In this work, we consider the problem of weakly-supervised multi-step localization in instructional videos. An established approach to this problem is to rely on a given list of steps. However, in reality, there is often more than one way to execute a procedure successfully, by following the set of…

2022

P3IV: Probabilistic Procedure Planning From Instructional Videos With Weak Supervision

CVPR 2022oral

In this paper, we study the problem of procedure planning in instructional videos. Here, an agent must produce a plausible sequence of actions that can transform the environment from a given start to a desired goal state. When learning procedure planning from instructional videos, most recent work l…

Cited by 52PDFcodeScholar
2021

Drop-DTW: Aligning Common Signal Between Sequences While Dropping Outliers

NeurIPS 2021poster

In this work, we consider the problem of sequence-to-sequence alignment for signals containing outliers. Assuming the absence of outliers, the standard Dynamic Time Warping (DTW) algorithm efficiently computes the optimal alignment between two (generally) variable-length sequences. While DTW is ro…

Cited by 59SourcePDFScholar
2020

Selecting Relevant Features from a Multi-domain Representation for Few-shot Classification

ECCV 2020poster

Popular approaches for few-shot classification consist of first learning a generic data representation based on a large annotated dataset, before adapting the representation to new classes given only a few labeled samples. In this work, we propose a new strategy based on feature selection, which is…

2018

Modeling Visual Context is Key to Augmenting Object Detection Datasets

ECCV 2018poster

Performing data augmentation for learning deep neural networks is well known to be important for training visual recognition systems. By artificially increasing the number of training examples, it helps reducing overfitting and improves generalization. For object detection, classical approaches for…

Cited by 314SourcePDFScholar
2017

BlitzNet: A Real-Time Deep Network for Scene Understanding

ICCV 2017poster

Real-time scene understanding has become crucial in many applications such as autonomous driving. In this paper, we propose a deep architecture, called BlitzNet, that jointly performs object detection and semantic segmentation in one forward pass, allowing real-time computations. Besides the computa…

Cited by 264PDFScholar