← Search

Vicky Kalogeiton

16 accepted papers

2026

IdEst: Assessing Self-Supervised Learning Representations via Intrinsic Dimension

ICML 2026poster

Self-supervised learning (SSL) has emerged as a powerful paradigm for learning meaningful representations from unlabeled data. However, the standard protocol for evaluating these representations, linear probing, is computationally expensive, sensitive to hyperparameters, and provides limited insight…

Cited by 0SourceScholar
2026

MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency

ICML 2026poster

The default paradigm of post-training text-to-image generators includes post-hoc selection of generated images, and subsequent training with one reward model to align the generator to the reward, typically user preference. This discards informative data as well as optimizes only for a single reward,…

Cited by 0SourceScholar
2026

Pulp Motion: Framing-aware multimodal camera and human motion generation

ICLR 2026poster

Treating human motion and camera trajectory generation separately overlooks a core principle of cinematography: the tight interplay between actor performance and camera work in the screen space. In this paper, we are the first to cast this task as a text-conditioned joint generation, aiming to main…

Cited by 0SourceScholar
2025

AKiRa: Augmentation Kit on Rays for Optical Video Generation

CVPR 2025poster

Recent advances in text-conditioned video diffusion have greatly improved video quality. However, these methods offer limited or sometimes no control to users on camera aspects, including dynamic camera motion, zoom, distorted lens and focus shifts. These motion and optical aspects are crucial for a…

2025

Around the World in 80 Timesteps: A Generative Approach to Global Visual Geolocation

CVPR 2025poster

Global visual geolocation predicts where an image was captured on Earth. Since images vary in how precisely they can be localized, this task inherently involves a significant degree of ambiguity. However, existing approaches are deterministic and overlook this aspect. In this paper, we aim to close…

2025

Di[M]O: Distilling Masked Diffusion Models into One-step Generator

ICCV 2025poster

Masked Diffusion Models (MDMs) have emerged as a powerful generative modeling technique. Despite their remarkable results, they typically suffer from slow inference with several steps. In this paper, we propose Di\mathtt [M] O, a novel approach that distills masked diffusion models into a one-step g…

2025

T-REGS: Minimum Spanning Tree Regularization for Self-Supervised Learning

NeurIPS 2025spotlight

Self-supervised learning (SSL) has emerged as a powerful paradigm for learning representations without labeled data, often by enforcing invariance to input transformations such as rotations or blurring. Recent studies have highlighted two pivotal properties for effective representations: (i) avoidin…

Cited by 0SourceScholar
2024

Collaborating Foundation Models for Domain Generalized Semantic Segmentation

CVPR 2024poster

Domain Generalized Semantic Segmentation (DGSS) deals with training a model on a labeled source domain with the aim of generalizing to unseen domains during inference. Existing DGSS methods typically effectuate robust features by means of Domain Randomization (DR). Such an approach is often limited…

2024

Don't Drop Your Samples! Coherence-Aware Training Benefits Conditional Diffusion

CVPR 2024highlight

Conditional diffusion models are powerful generative models that can leverage various types of conditional information such as class labels segmentation masks or text captions. However in many real-world scenarios conditional information may be noisy or unreliable due to human annotation errors or w…

Cited by 3SourcePDFScholar
2024

E.T. the Exceptional Trajectory: Text-to-camera-trajectory generation with character awareness

ECCV 2024poster

"Stories and emotions in movies emerge through the effect of well-thought-out directing decisions, in particular camera placement and movement over time. Crafting compelling camera trajectories remains a complex iterative process, even for skilful artists. To tackle this, in this paper, we propose a…

Cited by 3SourcePDFScholar
2022

SCAM! Transferring Humans between Images with Semantic Cross Attention Modulation

ECCV 2022poster

"A large body of recent work targets semantically conditioned image generation. Most such methods focus on the narrower task of pose transfer and ignore the more challenging task of subject transfer that consists in not only transferring the pose but also the appearance and background. In this work,…

2020

Smooth-AP: Smoothing the Path Towards Large-Scale Image Retrieval

ECCV 2020poster

Optimising a ranking-based metric, such as Average Precision (AP), is notoriously challenging due to the fact that it is non-differentiable, and hence cannot be optimised directly using gradient-descent methods. To this end, we introduce an objective that optimises instead a smoothed approximation o…

2019

LAEO-Net: Revisiting People Looking at Each Other in Videos

CVPR 2019poster

Capturing the 'mutual gaze' of people is essential for understanding and interpreting the social interactions between them. To this end, this paper addresses the problem of detecting people Looking At Each Other (LAEO) in video sequences. For this purpose, we propose LAEO-Net, a new deep CNN for det…

Cited by 75PDFcodeScholar
2017

Action Tubelet Detector for Spatio-Temporal Action Localization

ICCV 2017poster

Current state-of-the-art approaches for spatio-temporal action localization rely on detections at the frame level that are then linked or tracked across time. In this paper, we leverage the temporal continuity of videos instead of operating at the frame level. We propose the ACtion Tubelet detector…

Cited by 437PDFcodeScholar