← Search

Matej Kristan

22 accepted papers

2026

Mitigating Objectness Bias and Region-to-Text Misalignment for Open-Vocabulary Panoptic Segmentation

CVPR 2026

Open-vocabulary panoptic segmentation remains hindered by two coupled issues: (i) mask selection bias, where objectness heads trained on closed vocabularies suppress masks of categories not observed in training, and (ii) limited regional understanding in vision-language models such as CLIP, which we

Cited by 0SourcecodeScholar
2025

A Distractor-Aware Memory for Visual Object Tracking with SAM2

CVPR 2025poster

Memory-based trackers are video object segmentation methods that form the target model by concatenating recently tracked frames into a memory buffer and localize the target by attending the current image to the buffered frames. While already achieving top performance on many benchmarks, it was the r…

2024

A Novel Unified Architecture for Low-Shot Counting by Detection and Segmentation

NeurIPS 2024poster

Low-shot object counters estimate the number of objects in an image using few or no annotated exemplars. Objects are localized by matching them to prototypes, which are constructed by unsupervised image-wide object appearance aggregation. Due to potentially diverse object appearances, the existing a…

2024

Anomalous Sound Detection by Feature-Level Anomaly Simulation

ICASSP 2024accepted

Recently a growing number of works focus on machine defect detection from anomalous audio patterns. The datasets for the machine audio domain are scarce and recent methods that perform well on benchmarks such as DCASE2020 Task 2, rely on auxiliary information such as annotated data from other traini…

Cited by 0SourceScholar
2024

DAVE - A Detect-and-Verify Paradigm for Low-Shot Counting

CVPR 2024poster

Low-shot counters estimate the number of objects corresponding to a selected category based on only few or no exemplars annotated in the image. The current state-of-the-art estimates the total counts as the sum over the object location density map but do not provide object locations and sizes which…

2023

A Low-Shot Object Counting Network With Iterative Prototype Adaptation

ICCV 2023poster

We consider low-shot counting of arbitrary semantic categories in the image using only few annotated exemplars (few-shot) or no exemplars (no-shot). The standard few-shot pipeline follows extraction of appearance queries from exemplars and matching them with image features to infer the object counts…

Cited by 42PDFcodeScholar
2023

LaRS: A Diverse Panoptic Maritime Obstacle Detection Dataset and Benchmark

ICCV 2023poster

The progress in maritime obstacle detection is hindered by the lack of a diverse dataset that adequately captures the complexity of general maritime environments. We present the first maritime panoptic obstacle detection benchmark LaRS, featuring scenes from Lakes, Rivers and Seas. Our major contrib…

Cited by 28PDFcodeScholar
2022

DSR – A Dual Subspace Re-Projection Network for Surface Anomaly Detection

ECCV 2022poster

"The state-of-the-art in discriminative unsupervised surface anomaly detection relies on external datasets for synthesizing anomaly-augmented training images. Such approaches are prone to failure on near-in-distribution anomalies since these are difficult to be synthesized realistically due to their…

2021

DRAEM - A Discriminatively Trained Reconstruction Embedding for Surface Anomaly Detection

ICCV 2021poster

Visual surface anomaly detection aims to detect local image regions that significantly deviate from normal appearance. Recent surface anomaly detection methods rely on generative models to accurately reconstruct the normal areas and to fail on anomalies. These methods are trained only on anomaly-fre…

Cited by 820PDFcodeScholar
2019

CDTB: A Color and Depth Visual Object Tracking Dataset and Benchmark

ICCV 2019poster

We propose a new color-and-depth general visual object tracking benchmark (CDTB). CDTB is recorded by several passive and active RGB-D setups and contains indoor as well as outdoor sequences acquired in direct sunlight. The CDTB dataset is the largest and most diverse dataset in RGB-D tracking, with…

Cited by 89PDFcodeScholar
2019

Object Tracking by Reconstruction With View-Specific Discriminative Correlation Filters

CVPR 2019poster

Standard RGB-D trackers treat the target as a 2D structure, which makes modelling appearance changes related even to out-of-plane rotation challenging. This limitation is addressed by the proposed long-term RGB-D tracker called OTR - Object Tracking by Reconstruction. OTR performs online 3D target r…

Cited by 106PDFScholar
2019

The MaSTr1325 dataset for training deep USV obstacle detection models

IROS 2019poster

The progress of obstacle detection via semantic segmentation on unmanned surface vehicles (USVs) has been significantly lagging behind the developments in the related field of autonomous cars. The reason is the lack of large curated training datasets from USV domain required for development of data-…

Cited by 127SourceScholar
2017

Beyond Standard Benchmarks: Parameterizing Performance Evaluation in Visual Object Tracking

ICCV 2017poster

Object-to-camera motion produces a variety of apparent motion patterns that significantly affect performance of short-term visual trackers. Despite being crucial for designing robust trackers, their influence is poorly explored in standard benchmarks due to weakly defined, biased and overlapping att…

Cited by 26PDFScholar
2017

Discriminative Correlation Filter With Channel and Spatial Reliability

CVPR 2017poster

Short-term tracking is an open and challenging problem for which discriminative correlation filters (DCF) have shown excellent performance. We introduce the channel and spatial reliability concepts to DCF tracking and provide a novel learning algorithm for its efficient and seamless integration in…

Cited by 1672PDFcodeScholar
2016

Hierarchical spatial model for 2D range data based room categorization

ICRA 2016

The next generation service robots are expected to co-exist with humans in their homes. Such a mobile robot requires an efficient representation of space, which should be compact and expressive, for effective operation in real-world environments. In this paper we present a novel approach for 2D grou

Cited by 5SourceScholar