← Search

Vladimir Pavlovic

23 accepted papers

2025

GenVP: Generating Visual Puzzles with Contrastive Hierarchical VAEs

ICLR 2025poster

Raven’s Progressive Matrices (RPMs) is an established benchmark to examine the ability to perform high-level abstract visual reasoning (AVR). Despite the current success of algorithms that solve this task, humans can generalize beyond a given puzzle and create new puzzles given a set of rules, where…

Cited by 0SourcePDFScholar
2025

Hallucinatory Image Tokens: A Training-free EAZY Approach to Detecting and Mitigating Object Hallucinations in LVLMs

ICCV 2025poster

Despite their remarkable potential, Large Vision-Language Models (LVLMs) still face challenges with object hallucination, a problem where their generated outputs mistakenly incorporate objects that do not actually exist. Although most works focus on addressing this issue within the language-model ba…

Cited by 0SourcePDFScholar
2024

Learning from Synthetic Human Group Activities

CVPR 2024poster

The study of complex human interactions and group activities has become a focal point in human-centric computer vision. However progress in related tasks is often hindered by the challenges of obtaining large-scale labeled datasets from real-world scenarios. To address the limitation we introduce M3…

2023

ALWOD: Active Learning for Weakly-Supervised Object Detection

ICCV 2023poster

Object detection (OD), a crucial vision task, remains challenged by the lack of large training datasets with precise object localization labels. In this work, we propose ALWOD, a new framework that addresses this problem by fusing active learning (AL) with weakly and semi-supervised object detection…

Cited by 22PDFcodeScholar
2023

MSI: Maximize Support-Set Information for Few-Shot Segmentation

ICCV 2023poster

FSS (Few-shot segmentation) aims to segment a target class using a small number of labeled images (support set). To extract the information relevant to target class, a dominant approach in best performing FSS methods removes background features using a support mask. We observe that this feature exci…

Cited by 32PDFcodeScholar
2023

NP-SemiSeg: When Neural Processes meet Semi-Supervised Semantic Segmentation

ICML 2023poster

Semi-supervised semantic segmentation involves assigning pixel-wise labels to unlabeled images at training time. This is useful in a wide range of real-world applications where collecting pixel-wise labels is not feasible in time or cost. Current approaches to semi-supervised semantic segmentation w…

2022

Cross-Modal Coherence for Text-to-Image Retrieval

AAAI 2022technical

Common image-text joint understanding techniques presume that images and the associated text can universally be characterized by a single implicit model. However, co-occurring images and text can be related in qualitatively different ways, and explicitly modeling it could improve the performance of…

2022

HM: Hybrid Masking for Few-Shot Segmentation

ECCV 2022poster

"We study few-shot semantic segmentation that aims to segment a target object from a query image when provided with a few annotated support images of the target class. Several recent methods resort to a feature masking (FM) technique to discard irrelevant feature activations which eventually facilit…

2022

Harnessing Fourier Isovists and Geodesic Interaction for Long-Term Crowd Flow Prediction

IJCAI 2022poster

With the rise in popularity of short-term Human Trajectory Prediction (HTP), Long-Term Crowd Flow Prediction (LTCFP) has been proposed to forecast crowd movement in large and complex environments. However, the input representations, models, and datasets for LTCFP are currently limited. To this end,…

2022

MUSE-VAE: Multi-Scale VAE for Environment-Aware Long Term Trajectory Prediction

CVPR 2022poster

Accurate long-term trajectory prediction in complex scenes, where multiple agents (e.g., pedestrians or vehicles) interact with each other and the environment while attempting to accomplish diverse and often unknown goals, is a challenging stochastic forecasting problem. In this work, we propose MUS…

Cited by 90PDFScholar
2022

NP-Match: When Neural Processes meet Semi-Supervised Learning

ICML 2022spotlight

Semi-supervised learning (SSL) has been widely explored in recent years, and it is an effective way of leveraging unlabeled data to reduce the reliance on labeled data. In this work, we adjust neural processes (NPs) to the semi-supervised image classification task, resulting in a new method named NP…

2022

Variational Continual Proxy-Anchor for Deep Metric Learning

AISTATS 2022poster

The recent proxy-anchor method achieved outstanding performance in deep metric learning, which can be acknowledged to its data efficient loss based on hard example mining, as well as far lower sampling complexity than pair-based approaches. In this paper we extend the proxy-anchor method by posing i…

Cited by 2SourcePDFScholar
2021

CHEF: Cross-modal Hierarchical Embeddings for Food Domain Retrieval

AAAI 2021technical

Despite the abundance of multi-modal data, such as image-text pairs, there has been little effort in understanding the individual entities and their different roles in the construction of these data instances. In this work, we endeavour to discover the entities and their corresponding importance in…

2020

Laying the Foundations of Deep Long-Term Crowd Flow Prediction

ECCV 2020poster

Predicting the crowd behavior in complex environments is a key requirement for crowd and disaster management, architectural design, and urban planning. Given a crowd's immediate state, current approaches must be successively repeated over multiple time-steps for long-term predictions, leading to com…

2019

Bayes-Factor-VAE: Hierarchical Bayesian Deep Auto-Encoder Models for Factor Disentanglement

ICCV 2019oral

We propose a family of novel hierarchical Bayesian deep auto-encoder models capable of identifying disentangled factors of variability in data. While many recent attempts at factor disentanglement have focused on sophisticated learning objectives within the VAE framework, their choice of a standard…

Cited by 33PDFScholar
2019

Unsupervised Visual Domain Adaptation: A Deep Max-Margin Gaussian Process Approach

CVPR 2019oral

For unsupervised domain adaptation, the target domain error can be provably reduced by having a shared input representation that makes the source and target domains indistinguishable from each other. Very recently it has been shown that it is not only critical to match the marginal input distributio…

Cited by 54PDFScholar
2017

A Generative Model for Depth-Based Robust 3D Facial Pose Tracking

CVPR 2017poster

We consider the problem of depth-based robust 3D facial pose tracking under unconstrained scenarios with heavy occlusions and arbitrary facial expression variations. Unlike the previous depth-based discriminative or data-driven methods that require sophisticated training or manual intervention, we p…

Cited by 22PDFScholar
2017

Deep Structured Learning for Facial Action Unit Intensity Estimation

CVPR 2017poster

We consider the task of automated estimation of facial expression intensity. This involves estimation of multiple output variables (facial action units --- AUs) that are structurally dependent. Their structure arises from statistically induced co-occurrence patterns of AU intensity levels. Modeling…

Cited by 155PDFScholar
2017

PUnDA: Probabilistic Unsupervised Domain Adaptation for Knowledge Transfer Across Visual Categories

ICCV 2017poster

This paper introduces a probabilistic latent variable model to address unsupervised domain adaptation problems. This is achieved by learning projections from each domain to a latent space along the classifier in the latent space to simultaneously minimizing a notion of domain disparity while maximiz…

Cited by 46PDFScholar
2016

Copula Ordinal Regression for Joint Estimation of Facial Action Unit Intensity

CVPR 2016poster

Joint modeling of the intensity of facial action units (AUs) from face images is challenging due to the large number of AUs (30+) and their intensity levels (6). This is in part due to the lack of suitable models that can efficiently handle such a large number of outputs/classes simultaneously, but…

Cited by 72PDFScholar