← Search

Marius Leordeanu

11 accepted papers

2024

“Vorbești Românește?” A Recipe to Train Powerful Romanian LLMs with English Instructions

EMNLP 2024finding

In recent years, Large Language Models (LLMs) have achieved almost human-like performance on various tasks. While some LLMs have been trained on multilingual data, most of the training data is in English; hence, their performance in English greatly exceeds other languages. To our knowledge, we are t…

2022

UFO Depth: Unsupervised learning with flow-based odometry optimization for metric depth estimation

ICRA 2022poster

We propose an efficient method for unsupervised learning of metric depth estimation from a single image in the context of unconstrained videos captured from UAVs. We combine the accuracy of an analytical solution based on odometry with the power of deep learning. First, we show how to correct the no…

Cited by 6SourceScholar
2021

Discovering Dynamic Salient Regions for Spatio-Temporal Graph Neural Networks

NeurIPS 2021poster

Graph Neural Networks are perfectly suited to capture latent interactions between various entities in the spatio-temporal domain (e.g. videos). However, when an explicit structure is not available, it is not obvious what atomic elements should be represented as nodes. Current works generally use pre…

2021

Semi-Supervised Learning for Multi-Task Scene Understanding by Neural Graph Consensus

AAAI 2021technical

We address the challenging problem of semi-supervised learning in the context of multiple visual interpretations of the world by finding consensus in a graph of neural networks. Each graph node is a scene interpretation layer, while each edge is a deep net that transforms one layer at one node into…

Cited by 11SourcePDFScholar
2021

TeachText: CrossModal Generalized Distillation for Text-Video Retrieval

ICCV 2021poster

In recent years, considerable progress on the task of text-video retrieval has been achieved by leveraging large-scale pretraining on visual and audio datasets to construct powerful video encoders. By contrast, despite the natural symmetry, the design of effective algorithms for exploiting large-sca…

Cited by 167PDFcodeScholar
2020

A 3D Convolutional Approach to Spectral Object Segmentation in Space and Time

IJCAI 2020poster

We formulate object segmentation in video as a spectral graph clustering problem in space and time, in which nodes are pixels and their relations form local neighbourhoods. We claim that the strongest cluster in this pixel-level graph represents the salient object segmentation. We compute the main c…

2020

A hierarchical approach to vision-based language generation: from simple sentences to complex natural language

COLING 2020main

Automatically describing videos in natural language is an ambitious problem, which could bridge our understanding of vision and language. We propose a hierarchical approach, by first generating video descriptions as sequences of simple sentences, followed at the next level by a more complex and flue…

Cited by 7SourcePDFScholar
2017

Unsupervised Learning From Video to Detect Foreground Objects in Single Images

ICCV 2017poster

Unsupervised learning from visual data is one of the most difficult challenges in computer vision. It is essential for understanding how visual recognition works. Learning from unsupervised input has an immense practical value, as huge quantities of unlabeled videos can be collected at low cost. Her…

Cited by 65PDFScholar
2017

Unsupervised Object Segmentation in Video by Efficient Selection of Highly Probable Positive Features

ICCV 2017poster

We address an essential problem in computer vision, that of unsupervised foreground object segmentation in video, where a main object of interest in a video sequence should be automatically separated from its background. An efficient solution to this task would enable large-scale video interpretatio…

Cited by 38PDFScholar
2016

How Hard Can It Be? Estimating the Difficulty of Visual Search in an Image

CVPR 2016poster

We address the problem of estimating image difficulty defined as the human response time for solving a visual search task. We collect human annotations of image difficulty for the PASCAL VOC 2012 data set through a crowd-sourcing platform. We then analyze what human interpretable image properties ca…

Cited by 164PDFScholar