← Search

Efstratios Gavves

66 accepted papers

2026

Probing Newtonian Mechanics in Video Generative Models with Real Physical Systems

ICML 2026poster

Recent advances in image and video generation raise hopes that these models possess world modeling capabilities—the ability to generate realistic, physically plausible videos. This could revolutionize applications in robotics, autonomous driving, and scientific simulation. However, before treating t…

Cited by 0SourceScholar
2026

Towards Uniformity and Alignment for Multimodal Representation Learning

ICML 2026poster

Multimodal representation learning aims to construct a shared embedding space in which heterogeneous modalities are semantically aligned. Despite strong empirical results, InfoNCE-based objectives introduce inherent conflicts that yield distribution gaps across modalities. In this work, we identify …

Cited by 0SourceScholar
2025

Any-Resolution AI-Generated Image Detection by Spectral Learning

CVPR 2025poster

Recent works have established that AI models introduce spectral artifacts into generated images and propose approaches for learning to capture them using labeled data. However, the significant differences in such artifacts among different generative models hinder these approaches from generalizing t…

2025

CTRL-O: Language-Controllable Object-Centric Visual Representation Learning

CVPR 2025poster

Object-centric representation learning aims to decompose visual scenes into fixed-size vectors called "slots" or "object files", where each slot captures a distinct object. Current state-of-the-art object-centric models have shown remarkable success in object discovery in diverse domains including c…

Cited by 2SourcePDFScholar
2025

CaPo: Cooperative Plan Optimization for Efficient Embodied Multi-Agent Cooperation

ICLR 2025poster

In this work, we address the cooperation problem among large language model (LLM) based embodied agents, where agents must cooperate to achieve a common goal. Previous methods often execute actions extemporaneously and incoherently, without long-term strategic and cooperative planning, leading to r…

2025

Dream to Manipulate: Compositional World Models Empowering Robot Imitation Learning with Imagination

ICLR 2025poster

A world model provides an agent with a representation of its environment, enabling it to predict the causal consequences of its actions. Current world models typically cannot directly and explicitly imitate the actual environment in front of a robot, often resulting in unrealistic behaviors and hall…

2025

Grounding Continuous Representations in Geometry: Equivariant Neural Fields

ICLR 2025poster

Conditional Neural Fields (CNFs) are increasingly being leveraged as continuous signal representations, by associating each data-sample with a latent variable that conditions a shared backbone Neural Field (NeF) to reconstruct the sample. However, existing CNF architectures face limitations when usi…

2025

Language Agents Meet Causality -- Bridging LLMs and Causal World Models

ICLR 2025poster

Large Language Models (LLMs) have recently shown great promise in planning and reasoning applications. These tasks demand robust systems, which arguably require a causal understanding of the environment. While LLMs can acquire and reflect common sense causal knowledge from their pretraining data, th…

2025

LoTUS: Large-Scale Machine Unlearning with a Taste of Uncertainty

CVPR 2025poster

We present LoTUS, a novel Machine Unlearning (MU) method that eliminates the influence of training samples from pre-trained models, avoiding retraining from scratch. LoTUS smooths the prediction probabilities of the model up to an information-theoretic bound, mitigating its over-confidence stemming…

2025

Mechanistic PDE Networks for Discovery of Governing Equations

ICML 2025poster

We present Mechanistic PDE Networks -- a model for discovery of governing *partial differential equations* from data. Mechanistic PDE Networks represent spatiotemporal data as space-time dependent *linear* partial differential equations in neural network hidden representations. The represented PDEs…

Cited by 1SourcePDFScholar
2025

MoSiC: Optimal-Transport Motion Trajectory for Dense Self-Supervised Learning

ICCV 2025poster

Dense self-supervised learning has shown great promise for learning pixel- and patch-level representations, but extending it to videos remains challenging due to the complexity of motion dynamics. Existing approaches struggle as they rely on static augmentations that fail under object deformations,…

2025

Probabilistic Interactive 3D Segmentation with Hierarchical Neural Processes

ICML 2025poster

Interactive 3D segmentation has emerged as a promising solution for generating accurate object masks in complex 3D scenes by incorporating user-provided clicks. However, two critical challenges remain underexplored: (1) effectively generalizing from sparse user clicks to produce accurate segmentatio…

Cited by 0SourcePDFScholar
2025

Probabilistic Prototype Calibration of Vision-language Models for Generalized Few-shot Semantic Segmentation

ICCV 2025poster

Generalized Few-Shot Semantic Segmentation (GFSS) aims to extend a segmentation model to novel classes with only a few annotated examples while maintaining performance on base classes. Recently, pretrained vision-language models (VLMs) such as CLIP have been leveraged in GFSS to improve generalizati…

2024

Click Prompt Learning with Optimal Transport for Interactive Segmentation

ECCV 2024poster

"Click-based interactive segmentation aims to segment target objects conditioned on user-provided clicks. Existing methods typically interpret user intention by learning multiple click prompts to generate corresponding prompt-activated masks, and selecting one from these masks. However, directly mat…

2024

Domain Adaptation with Cauchy-Schwarz Divergence

UAI 2024poster

Domain adaptation aims to use training data from one or multiple source domains to learn a hypothesis that can be generalized to a different, but related, target domain. As such, having a reliable measure for evaluating the discrepancy of both marginal and conditional distributions is crucial. We in…

2024

GeneralAD: Anomaly Detection Across Domains by Attending to Distorted Features

ECCV 2024poster

"In the domain of anomaly detection, methods often excel in either high-level semantic or low-level industrial benchmarks, rarely achieving cross-domain proficiency. Semantic anomalies are novelties that differ in meaning from the training set, like unseen objects in self-driving cars. In contrast,…

2024

Graph Neural Networks for Learning Equivariant Representations of Neural Networks

ICLR 2024oral

Neural networks that process the parameters of other neural networks find applications in domains as diverse as classifying implicit neural representations, generating neural network weights, and predicting generalization errors. However, existing approaches either overlook the inherent permutation…

2024

How to Train Neural Field Representations: A Comprehensive Study and Benchmark

CVPR 2024poster

Neural fields (NeFs) have recently emerged as a versatile method for modeling signals of various modalities including images shapes and scenes. Subsequently a number of works have explored the use of NeFs as representations for downstream tasks e.g. classifying an image based on the parameters of a…

2024

SIGMA: Sinkhorn-Guided Masked Video Modeling

ECCV 2024poster

"Video-based pretraining offers immense potential for learning strong visual representations on an unprecedented scale. Recently, masked video modeling methods have shown promising scalability, yet fall short in capturing higher-level semantics due to reconstructing predefined low-level targets such…

Cited by 2SourcePDFScholar
2024

VISA: Reasoning Video Object Segmentation via Large Language Model

ECCV 2024poster

"Existing Video Object Segmentation (VOS) relies on explicit user instructions, such as categories, masks, or short phrases, restricting their ability to perform complex video segmentation requiring reasoning with world knowledge. In this paper, we introduce a new task, Reasoning Video Object Segmen…

2023

BISCUIT: Causal Representation Learning from Binary Interactions

UAI 2023poster

Identifying the causal variables of an environment and how to intervene on them is of core value in applications such as robotics and embodied AI. While an agent can commonly interact with the environment and may implicitly perturb the behavior of some of these causal variables, often the targets it…

2023

Causal Representation Learning for Instantaneous and Temporal Effects in Interactive Systems

ICLR 2023poster

Causal representation learning is the task of identifying the underlying causal variables and their relations from high-dimensional observations, such as images. Recent work has shown that one can reconstruct the causal variables from temporal sequences of observations under the assumption that ther…

2023

Differentiable Mathematical Programming for Object-Centric Representation Learning

ICLR 2023poster

We propose topology-aware feature partitioning into $k$ disjoint partitions for given scene features as a method for object-centric representation learning. To this end, we propose to use minimum $s$-$t$ graph cuts as a partitioning method which is represented as a linear program. The method is topo…

Cited by 7SourcePDFScholar
2023

Fake It Until You Make It : Towards Accurate Near-Distribution Novelty Detection

ICLR 2023poster

We aim for image-based novelty detection. Despite considerable progress, existing models either fail or face dramatic drop under the so-called ``near-distribution" setup, where the differences between normal and anomalous samples are subtle. We first demonstrate existing methods could experience up…

Cited by 34SourcePDFScholar
2023

Latent Field Discovery in Interacting Dynamical Systems with Neural Fields

NeurIPS 2023poster

Systems of interacting objects often evolve under the influence of underlying field effects that govern their dynamics, yet previous works have abstracted away from such effects, and assume that systems evolve in a vacuum. In this work, we focus on discovering these fields, and infer them from the o…

2023

Modelling Long Range Dependencies in $N$D: From Task-Specific to a General Purpose CNN

ICLR 2023poster

Performant Convolutional Neural Network (CNN) architectures must be tailored to specific tasks in order to consider the length, resolution, and dimensionality of the input data. In this work, we tackle the need for problem-specific CNN architectures. We present the Continuous Convolutional Neural Ne…

2023

Modulated Neural ODEs

NeurIPS 2023poster

Neural ordinary differential equations (NODEs) have been proven useful for learning non-linear dynamics of arbitrary trajectories. However, current NODE methods capture variations across trajectories only via the initial state value or by auto-regressive encoder updates. In this work, we introduce M…

2023

Time Does Tell: Self-Supervised Time-Tuning of Dense Image Representations

ICCV 2023poster

Spatially dense self-supervised learning is a rapidly growing problem domain with promising applications for unsupervised segmentation and pretraining for dense downstream tasks. Despite the abundance of temporal data in the form of videos, this information-rich source has been largely overlooked. O…

Cited by 21PDFcodeScholar
2023

Towards Open-Vocabulary Video Instance Segmentation

ICCV 2023oral

Video Instance Segmentation (VIS) aims at segmenting and categorizing objects in videos from a closed set of training categories, lacking the generalization ability to handle novel categories in real-world videos. To address this limitation, we make the following three contributions. First, we intro…

Cited by 36PDFcodeScholar
2022

3D Equivariant Graph Implicit Functions

ECCV 2022poster

"In recent years, neural implicit representations have made remarkable progress in modeling of 3D shapes with arbitrary topology. In this work, we address two key limitations of such representations, in failing to capture local 3D geometric fine details, and to learn from and generalize to shapes wi…

2022

Batch Bayesian Optimization on Permutations using the Acquisition Weighted Kernel

NeurIPS 2022accept

In this work we propose a batch Bayesian optimization method for combinatorial problems on permutations, which is well suited for expensive-to-evaluate objectives. We first introduce LAW, an efficient batch acquisition method based on determinantal point processes using the acquisition weighted kern…

2022

Delta Distillation for Efficient Video Processing

ECCV 2022poster

"This paper aims to accelerate video stream processing, such as object detection and semantic segmentation, by leveraging the temporal redundancies that exist between video frames. Instead of relying on explicit motion alignment, such as optical flow warping, we propose a novel knowledge distillatio…

2022

Dynamic Prototype Convolution Network for Few-Shot Semantic Segmentation

CVPR 2022poster

The key challenge for few-shot semantic segmentation (FSS) is how to tailor a desirable interaction among support and query features and/or their prototypes, under the episodic training scenario. Most existing FSS methods implement such support/query interactions by solely leveraging \it plain oper…

Cited by 111PDFScholar
2022

NFormer: Robust Person Re-Identification With Neighbor Transformer

CVPR 2022poster

Person re-identification aims to retrieve persons in highly varying settings across different cameras and scenarios, in which robust and discriminative representation learning is crucial. Most research considers learning representations from single images, ignoring any potential interactions between…

Cited by 175PDFcodeScholar
2021

Mixed variable Bayesian optimization with frequency modulated kernels

UAI 2021poster

The sample efficiency of Bayesian optimization(BO) is often boosted by Gaussian Process(GP) surrogate models. However, on mixed variable spaces, surrogate models other than GPs are prevalent, mainly due to the lack of kernels which can model complex dependencies across different types of variables.…

2021

Neural Feature Matching in Implicit 3D Representations

ICML 2021spotlight

Recently, neural implicit functions have achieved impressive results for encoding 3D shapes. Conditioning on low-dimensional latent codes generalises a single implicit function to learn shared representation space for a variety of shapes, with the advantage of smooth interpolation. While the benefit…

2021

Roto-translated Local Coordinate Frames For Interacting Dynamical Systems

NeurIPS 2021poster

Modelling interactions is critical in learning complex dynamical systems, namely systems of interacting objects with highly non-linear and time-dependent behaviour. A large class of such systems can be formalized as $\textit{geometric graphs}$, $\textit{i.e.}$ graphs with nodes positioned in the Euc…

2021

Sparse-Shot Learning With Exclusive Cross-Entropy for Extremely Many Localisations

ICCV 2021poster

Object localisation, in the context of regular images, often depicts objects like people or cars. In these images, there is typically a relatively small number of objects per class, which usually is manageable to annotate. However, outside the setting of regular images, we are often confronted with…

Cited by 2PDFScholar
2021

Spectral Smoothing Unveils Phase Transitions in Hierarchical Variational Autoencoders

ICML 2021oral

Variational autoencoders with deep hierarchies of stochastic layers have been known to suffer from the problem of posterior collapse, where the top layers fall back to the prior and become independent of input. We suggest that the hierarchical VAE objective explicitly includes the variance of the fu…

Cited by 10SourcePDFScholar
2020

Low Bias Low Variance Gradient Estimates for Boolean Stochastic Networks

ICML 2020poster

Stochastic neural networks with discrete random variables are an important class of models for their expressiveness and interpretability. Since direct differentiation and backpropagation is not possible, Monte Carlo gradient estimation techniques are a popular alternative. Efficient stochastic gradi…

Cited by 17SourcePDFScholar
2020

PointMixup: Augmentation for Point Clouds

ECCV 2020poster

This paper introduces data augmentation for point clouds by interpolation between examples. Data augmentation by interpolation has shown to be a simple and effective approach in the image domain. Such a mixup is however not directly transferable to point clouds, as we do not have a one-to-one corres…

2019

Combinatorial Bayesian Optimization using the Graph Cartesian Product

NeurIPS 2019poster

This paper focuses on Bayesian Optimization (BO) for objectives on combinatorial search spaces, including ordinal and categorical variables. Despite the abundance of potential applications of Combinatorial BO, including chipset configuration search and neural architecture search, only a handful of m…

2019

Relaxed Quantization for Discretized Neural Networks

ICLR 2019poster

Neural network quantization has become an important research area due to its great impact on deployment of large models on resource constrained devices. In order to train networks that can be effectively discretized without loss of performance, we introduce a differentiable quantization procedure. D…

Cited by 224SourcePDFScholar
2019

Spherical Regression: Learning Viewpoints, Surface Normals and 3D Rotations on N-Spheres

CVPR 2019poster

Many computer vision challenges require continuous outputs, but tend to be solved by discrete classification. The reason is classification's natural containment within a probability n-simplex, as defined by the popular softmax activation function. Regular regression lacks such a closed geometry, lea…

Cited by 78PDFcodeScholar
2019

Training a Spiking Neural Network with Equilibrium Propagation

AISTATS 2019poster

Backpropagation is almost universally used to train artificial neural networks. However, there are several reasons that backpropagation could not be plausibly implemented by biological neurons. Among these are the facts that (1) biological neurons appear to lack any mechanism for sending gradients…

Cited by 54SourcePDFScholar
2018

Long-term Tracking in the Wild: a Benchmark

ECCV 2018poster

We introduce the OxUvA dataset and benchmark for evaluating single-object tracking algorithms. Benchmarks have enabled great strides in the field of object tracking by defining standardized evaluations on large sets of diverse videos. However, these works have focused exclusively on sequences that a…

Cited by 206SourcePDFScholar
2018

Temporally Efficient Deep Learning with Spikes

ICLR 2018poster

The vast majority of natural sensory data is temporally redundant. For instance, video frames or audio samples which are sampled at nearby points in time tend to have similar values. Typically, deep learning algorithms take no advantage of this redundancy to reduce computations. This can be an obs…

2017

Self-Supervised Video Representation Learning With Odd-One-Out Networks

CVPR 2017poster

We propose a new self-supervised CNN pre-training technique based on a novel auxiliary task called odd-one-out learning. In this task, the machine is asked to identify the unrelated or odd element from a set of otherwise related elements. We apply this technique to self-supervised video representati…

Cited by 562PDFScholar
2017

Tracking by Natural Language Specification

CVPR 2017poster

This paper strives to track a target object in a video. Rather than specifying the target in the first frame of a video by a bounding box, we propose to track the object based on a natural language specification of the target, which provides a more natural human-machine interaction as well as a mean…

Cited by 206PDFScholar
2017

Unified Embedding and Metric Learning for Zero-Exemplar Event Detection

CVPR 2017poster

Event detection in unconstrained videos is conceived as a content-based video retrieval with two modalities: textual and visual. Given a text describing a novel event, the goal is to rank related videos accordingly. This task is zero-exemplar, no video examples are given to the novel event. Related…

Cited by 19PDFcodeScholar
2016

Dynamic Image Networks for Action Recognition

CVPR 2016oral

We introduce the concept of dynamic image, a novel compact representation of videos useful for video analysis especially when convolutional neural networks (CNNs) are used. The dynamic image is based on the rank pooling concept and is obtained through the parameters of a ranking machine that encodes…

Cited by 727PDFcodeScholar
2015

Active Transfer Learning With Zero-Shot Priors: Reusing Past Datasets for Future Tasks

ICCV 2015poster

How can we reuse existing knowledge, in the form of available datasets, when solving a new and apparently unrelated target task from a set of unlabeled data? In this work we make a first contribution to answer this question in the context of image classification. We frame this quest as an active…

Cited by 88PDFScholar
2015

Guiding the Long-Short Term Memory Model for Image Caption Generation

ICCV 2015poster

In this work we focus on the problem of image caption generation. We propose an extension of the long short term memory (LSTM) model, which we coin gLSTM for short. In particular, we add semantic information extracted from the image as extra input to each unit of the LSTM block, with the aim of gui…

Cited by 587PDFcodeScholar
2015

Modeling Video Evolution for Action Recognition

CVPR 2015poster

In this paper we present a method to capture video-wide temporal information for action recognition. We postulate that a function capable of ordering the frames of a video temporally (based on the appearance) captures well the evolution of the appearance within the video. We learn such ranking funct…

Cited by 709SourcePDFScholar