← Search

Cees G. M. Snoek

75 accepted papers

2026

GateRA: Token-aware Modulation for Parameter-Efficient Fine-tuning

AAAI 2026technical

Parameter-efficient fine-tuning (PEFT) methods, such as LoRA, DoRA, and HiRA, enable lightweight adaptation of large pre-trained models via low-rank updates. However, existing PEFT approaches apply static, input-agnostic updates to all tokens, disregarding the varying importance and difficulty of d

Cited by 0SourcePDFScholar
2026

Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs

CVPR 2026

While continual visual instruction tuning (CVIT) has shown promise in adapting multimodal large language models (MLLMs), existing studies predominantly focus on models without safety alignment. This critical oversight ignores the fact that real-world MLLMs inherently require such mechanisms to mitig

Cited by 0SourcecodeScholar
2026

MoAlign: Motion-Centric Representation Alignment for Video Diffusion Models

ICLR 2026poster

Text-to-video diffusion models have enabled high-quality video synthesis, yet often fail to generate temporally coherent and physically plausible motion. A key reason is the models' insufficient understanding of complex motions that natural videos often entail. Recent works tackle this problem by al…

Cited by 0SourceScholar
2026

Prompt-Robust Vision-Language Models via Meta-Finetuning

ICLR 2026poster

Vision-language models (VLMs) have demonstrated remarkable generalization across diverse tasks by leveraging large-scale image-text pretraining. However, their performance is notoriously unstable under variations in natural language prompts, posing a considerable challenge for reliable real-world de…

Cited by 0SourceScholar
2026

Purrception: Variational Flow Matching for Vector-Quantized Image Generation

ICLR 2026poster

We introduce Purrception, a variational flow matching approach for vector-quantized image generation that provides explicit categorical supervision while maintaining continuous transport dynamics. Our method adapts Variational Flow Matching to vector-quantized latents by learning categorical posteri…

Cited by 0SourceScholar
2026

RegionReasoner: Region-Grounded Multi-Round Visual Reasoning

ICLR 2026poster

Large vision-language models have achieved remarkable progress in visual reasoning, yet most existing systems rely on single-step or text-only reasoning, limiting their ability to iteratively refine understanding across multiple visual contexts. To address this limitation, we introduce a new multi-r…

Cited by 0SourcecodeScholar
2026

What Layers When: Learning to Skip Compute in LLMs with Residual Gates

ICLR 2026poster

We introduce GateSkip, a simple residual-stream gating mechanism that enables token-wise layer skipping in decoder-only LMs. Each Attention/MLP branch is equipped with a sigmoid-linear gate that compresses the branch’s output before it re-enters the residual stream. During inference we rank tokens b…

Cited by 0SourcecodeScholar
2025

An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels

ICLR 2025poster

This work does not introduce a new method. Instead, we present an interesting finding that questions the necessity of the inductive bias of locality in modern computer vision architectures. Concretely, we find that vanilla Transformers can operate by directly treating each individual pixel as a toke…

Cited by 13SourcePDFScholar
2025

CaPo: Cooperative Plan Optimization for Efficient Embodied Multi-Agent Cooperation

ICLR 2025poster

In this work, we address the cooperation problem among large language model (LLM) based embodied agents, where agents must cooperate to achieve a common goal. Previous methods often execute actions extemporaneously and incoherently, without long-term strategic and cooperative planning, leading to r…

2025

Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning

CVPR 2025poster

This paper proposes the first video-grounded entailment tree reasoning method for commonsense video question answering (VQA). Despite the remarkable progress of large visual-language models (VLMs), there are growing concerns that they learn spurious correlations between videos and likely answers, re…

Cited by 0SourcePDFScholar
2025

DynaPrompt: Dynamic Test-Time Prompt Tuning

ICLR 2025poster

Test-time prompt tuning enhances zero-shot generalization of vision-language models but tends to ignore the relatedness among test samples during inference. Online test-time prompt tuning provides a simple way to leverage the information in previous test samples, albeit with the risk of prompt colla…

Cited by 0SourcePDFScholar
2025

Elastic ViTs from Pretrained Models without Retraining

NeurIPS 2025poster

Vision foundation models achieve remarkable performance but are only available in a limited set of pre-determined sizes, forcing sub-optimal deployment choices under real-world constraints. We introduce SnapViT: single-shot network approximation for pruned Vision Transformers, a new post-pretraining…

Cited by 0SourceScholar
2025

MoSiC: Optimal-Transport Motion Trajectory for Dense Self-Supervised Learning

ICCV 2025poster

Dense self-supervised learning has shown great promise for learning pixel- and patch-level representations, but extending it to videos remains challenging due to the complexity of motion dynamics. Existing approaches struggle as they rely on static augmentations that fail under object deformations,…

2025

One Hundred Neural Networks and Brains Watching Videos: Lessons from Alignment

ICLR 2025poster

What can we learn from comparing video models to human brains, arguably the most efficient and effective video processing systems in existence? Our work takes a step towards answering this question by performing the first large-scale benchmarking of deep video models on representational alignment to…

Cited by 0SourcePDFScholar
2025

TULIP: Token-length Upgraded CLIP

ICLR 2025poster

We address the challenge of representing long captions in vision-language models, such as CLIP. By design these models are limited by fixed, absolute positional encodings, restricting inputs to a maximum of 77 tokens and hindering performance on tasks requiring longer descriptions. Although recent w…

2025

TWIST & SCOUT: Grounding Multimodal LLM-Experts by Forget-Free Tuning

ICCV 2025poster

Spatial awareness is key to enable embodied multimodal AI systems. Yet, without vast amounts of spatial supervision, current Multimodal Large Language Models (MLLMs) struggle at this task. In this paper, we introduce TWIST & SCOUT, a framework that equips pre-trained MLLMs with visual grounding abil…

Cited by 0SourcePDFScholar
2025

The Sound of Water: Inferring Physical Properties from Pouring Liquids

ICASSP 2025accepted

We study the connection between audio-visual observations and the underlying physics of a mundane yet intriguing everyday activity: pouring liquids. Given only the sound of liquid pouring into a container, our objective is to automatically infer physical properties such as the liquid level, the shap…

Cited by 0SourceScholar
2025

Union-over-Intersections: Object Detection beyond Winner-Takes-All

ICLR 2025spotlight

This paper revisits the problem of predicting box locations in object detection architectures. Typically, each box proposal or box query aims to directly maximize the intersection-over-union score with the ground truth, followed by a winner-takes-all non-maximum suppression where only the highest sc…

2024

Any-Shift Prompting for Generalization over Distributions

CVPR 2024poster

Image-language models with prompt learning have shown remarkable advances in numerous downstream vision tasks. Nevertheless conventional prompt learning methods overfit the training distribution and lose the generalization ability on the test distributions. To improve the generalization across vario…

Cited by 17SourcePDFScholar
2024

Graph Neural Networks for Learning Equivariant Representations of Neural Networks

ICLR 2024oral

Neural networks that process the parameters of other neural networks find applications in domains as diverse as classifying implicit neural representations, generating neural network weights, and predicting generalization errors. However, existing approaches either overlook the inherent permutation…

2024

IPO: Interpretable Prompt Optimization for Vision-Language Models

NeurIPS 2024poster

Pre-trained vision-language models like CLIP have remarkably adapted to various downstream tasks. Nonetheless, their performance heavily depends on the specificity of the input text prompts, which requires skillful prompt template engineering. Instead, current approaches to prompt optimization learn…

2024

PIN: Positional Insert Unlocks Object Localisation Abilities in VLMs

CVPR 2024poster

Vision-Language Models (VLMs) such as Flamingo and GPT-4V have shown immense potential by integrating large language models with vision systems. Nevertheless these models face challenges in the fundamental computer vision task of object localisation due to their training on multimodal data containin…

Cited by 13SourcePDFScholar
2024

R-MAE: Regions Meet Masked Autoencoders

ICLR 2024poster

In this work, we explore regions as a potential visual analogue of words for self-supervised image representation learning. Inspired by Masked Autoencoding (MAE), a generative pre-training baseline, we propose masked region autoencoding to learn from groups of pixels or regions. Specifically, we des…

2023

Detecting Objects with Context-Likelihood Graphs and Graph Refinement

ICCV 2023poster

The goal of this paper is to detect objects by exploiting their interrelationships. Contrary to existing methods, which learn objects and relations separately, our key idea is to learn the object-relation distribution jointly. We first propose a novel way of creating a graphical representation of an…

Cited by 2PDFScholar
2023

Energy-Based Test Sample Adaptation for Domain Generalization

ICLR 2023poster

In this paper, we propose energy-based sample adaptation at test time for domain generalization. Where previous works adapt their models to target domains, we adapt the unseen target samples to source-trained models. To this end, we design a discriminative energy-based model, which is trained on sou…

2023

Fake It Until You Make It : Towards Accurate Near-Distribution Novelty Detection

ICLR 2023poster

We aim for image-based novelty detection. Despite considerable progress, existing models either fail or face dramatic drop under the so-called ``near-distribution" setup, where the differences between normal and anomalous samples are subtle. We first demonstrate existing methods could experience up…

Cited by 34SourcePDFScholar
2023

Learn to Categorize or Categorize to Learn? Self-Coding for Generalized Category Discovery

NeurIPS 2023poster

In the quest for unveiling novel categories at test time, we confront the inherent limitations of traditional supervised recognition models that are restricted by a predefined category set. While strides have been made in the realms of self-supervised and open-world learning towards test-time catego…

2023

MetaModulation: Learning Variational Feature Hierarchies for Few-Shot Learning with Fewer Tasks

ICML 2023poster

Meta-learning algorithms are able to learn a new task using previously learned knowledge, but they often require a large number of meta-training tasks which may not be readily available. To address this issue, we propose a method for few-shot learning with fewer tasks, which we call MetaModulation.…

2023

Order-preserving Consistency Regularization for Domain Adaptation and Generalization

ICCV 2023poster

Deep learning models fail on cross-domain challenges if the model is oversensitive to domain-specific attributes, e.g., lightning, background, camera angle, etc. To alleviate this problem, data augmentation coupled with consistency regularization are commonly adopted to make the model less sensitive…

Cited by 16PDFcodeScholar
2023

ProtoDiff: Learning to Learn Prototypical Networks by Task-Guided Diffusion

NeurIPS 2023poster

Prototype-based meta-learning has emerged as a powerful technique for addressing few-shot learning challenges. However, estimating a deterministic prototype using a simple average function from a limited number of examples remains a fragile process. To overcome this limitation, we introduce ProtoDif…

2023

Self-Guided Diffusion Models

CVPR 2023poster

Diffusion models have demonstrated remarkable progress in image generation quality, especially when guidance is used to control the generative process. However, guidance requires a large amount of image-annotation pairs for training and is thus dependent on their availability and correctness. In thi…

2023

SuperDisco: Super-Class Discovery Improves Visual Recognition for the Long-Tail

CVPR 2023poster

Modern image classifiers perform well on populated classes while degrading considerably on tail classes with only a few instances. Humans, by contrast, effortlessly handle the long-tailed recognition challenge, since they can learn the tail representation based on different levels of semantic abstra…

Cited by 18SourcePDFScholar
2023

Test of Time: Instilling Video-Language Models With a Sense of Time

CVPR 2023poster

Modelling and understanding time remains a challenge in contemporary video understanding models. With language emerging as a key driver towards powerful generalization, it is imperative for foundational video-language models to have a sense of time. In this paper, we consider a specific aspect of te…

2023

Tubelet-Contrastive Self-Supervision for Video-Efficient Generalization

ICCV 2023poster

We propose a self-supervised method for learning motion-focused video representations. Existing approaches minimize distances between temporally augmented videos, which maintain high spatial similarity. We instead propose to learn similarities between videos with identical local motion dynamics but…

Cited by 13PDFcodeScholar
2023

Unlocking Slot Attention by Changing Optimal Transport Costs

ICML 2023poster

Slot attention is a powerful method for object-centric modeling in images and videos. However, its set-equivariance limits its ability to handle videos with a dynamic number of objects because it cannot break ties. To overcome this limitation, we first establish a connection between slot attention a…

2022

Association Graph Learning for Multi-Task Classification with Category Shifts

NeurIPS 2022accept

In this paper, we focus on multi-task classification, where related classification tasks share the same label space and are learned simultaneously. In particular, we tackle a new setting, which is more realistic than currently addressed in the literature, where categories shift from training to test…

2022

BoxeR: Box-Attention for 2D and 3D Transformers

CVPR 2022poster

In this paper, we propose a simple attention mechanism, we call Box-Attention. It enables spatial interaction between grid features, as sampled from boxes of interest, and improves the learning capability of transformers for several vision tasks. Specifically, we present BoxeR, short for Box Transfo…

Cited by 43PDFcodeScholar
2022

Hierarchical Variational Memory for Few-shot Learning Across Domains

ICLR 2022poster

Neural memory enables fast adaptation to new tasks with just a few training samples. Existing memory models store features only from the single last layer, which does not generalize well in presence of a domain shift between training and test distributions. Rather than relying on a flat memory, we p…

2022

How Severe Is Benchmark-Sensitivity in Video Self-Supervised Learning?

ECCV 2022poster

"Despite the recent success of video self-supervised learning models, there is much still to be understood about their generalization capability. In this paper, we investigate how sensitive video self-supervised learning is to the current conventional benchmark and whether methods generalize beyond…

2022

Learning to Generalize across Domains on Single Test Samples

ICLR 2022poster

We strive to learn a model from a set of source domains that generalizes well to unseen target domains. The main challenge in such a domain generalization scenario is the unavailability of any target domain data during training, resulting in the learned model not being explicitly adapted to the unse…

2022

Less than Few: Self-Shot Video Instance Segmentation

ECCV 2022poster

"The goal of this paper is to bypass the need for labelled examples in few-shot video understanding at run time. While proven effective, in many practical video settings even labelling a few examples appears unrealistic. This is especially true as the level of details in spatio-temporal video unders…

2022

Multiset-Equivariant Set Prediction with Approximate Implicit Differentiation

ICLR 2022poster

Most set prediction models in deep learning use set-equivariant operations, but they actually operate on multisets. We show that set-equivariant functions cannot represent certain functions on multisets, so we introduce the more appropriate notion of multiset-equivariance. We identify that the exist…

2022

TubeR: Tubelet Transformer for Video Action Detection

CVPR 2022oral

We propose TubeR: a simple solution for spatio-temporal video action detection. Different from existing methods that depend on either an off-line actor detector or hand-designed actor-positional hypotheses like proposals or anchors, we propose to directly detect an action tubelet in video by simulta…

Cited by 95PDFScholar
2022

Variational Model Perturbation for Source-Free Domain Adaptation

NeurIPS 2022accept

We aim for source-free domain adaptation, where the task is to deploy a model pre-trained on source domains to target domains. The challenges stem from the distribution shift from the source to the target domain, coupled with the unavailability of any source data and labeled target data for optimiza…

2021

MetaNorm: Learning to Normalize Few-Shot Batches Across Domains

ICLR 2021poster

Batch normalization plays a crucial role when training deep neural networks. However, batch statistics become unstable with small batch sizes and are unreliable in the presence of distribution shifts. We propose MetaNorm, a simple yet effective meta-learning normalization. It tackles the aforementio…

Cited by 83SourcePDFScholar
2021

Motion-Augmented Self-Training for Video Recognition at Smaller Scale

ICCV 2021poster

The goal of this paper is to self-train a 3D convolutional neural network on an unlabeled video collection for deployment on small-scale video collections. As smaller video datasets benefit more from motion than appearance, we strive to train our network using optical flow, but avoid its computation…

Cited by 24PDFScholar
2021

Set Prediction without Imposing Structure as Conditional Density Estimation

ICLR 2021poster

Set prediction is about learning to predict a collection of unordered variables with unknown interrelations. Training such models with set losses imposes the structure of a metric space over sets. We focus on stochastic and underdefined cases, where an incorrectly chosen loss function leads to impla…

2021

Social Fabric: Tubelet Compositions for Video Relation Detection

ICCV 2021poster

This paper strives to classify and detect the relationship between object tubelets appearing within a video as a <subject-predicate-object> triplet. Where existing works treat object proposals or tubelets as single entities and model their relations a posteriori, we propose to classify and detect pr…

Cited by 30PDFcodeScholar
2020

Cloth in the Wind: A Case Study of Physical Measurement Through Simulation

CVPR 2020poster

For many of the physical phenomena around us, we have developed sophisticated models explaining their behavior. Nevertheless, measuring physical properties from visual observations is challenging due to the high number of causally underlying physical parameters -- including material properties and e…

Cited by 36PDFScholar
2020

Latent Embedding Feedback and Discriminative Features for Zero-Shot Classification

ECCV 2020poster

Zero-shot learning strives to classify unseen categories for which no data is available during training. In the generalized variant, the test samples can further belong to seen or unseen categories. The state-of-the-art relies on Generative Adversarial Networks that synthesize unseen class features…

2020

Learning to Learn with Variational Information Bottleneck for Domain Generalization

ECCV 2020poster

Domain generalization models learn to generalize to previously unseen domains, but suffer from prediction uncertainty and domain shift. In this paper, we address both problems. We introduce a probabilistic meta-learning model for domain generalization, in which classifier parameters shared across do…

Cited by 196SourcePDFScholar
2020

Localizing the Common Action Among a Few Videos

ECCV 2020poster

This paper strives to localize the temporal extent of an action in a long untrimmed video. Where existing work leverages many examples with their start, their ending, and/or the class of the action during training time, we propose few-shot common action localization. The start and end of an action i…

2020

PointMixup: Augmentation for Point Clouds

ECCV 2020poster

This paper introduces data augmentation for point clouds by interpolation between examples. Data augmentation by interpolation has shown to be a simple and effective approach in the image domain. Such a mixup is however not directly transferable to point clouds, as we do not have a one-to-one corres…

2019

Spherical Regression: Learning Viewpoints, Surface Normals and 3D Rotations on N-Spheres

CVPR 2019poster

Many computer vision challenges require continuous outputs, but tend to be solved by discrete classification. The reason is classification's natural containment within a probability n-simplex, as defined by the popular softmax activation function. Regular regression lacks such a closed geometry, lea…

Cited by 78PDFcodeScholar
2018

Actor and Action Video Segmentation From a Sentence

CVPR 2018poster

This paper strives for pixel-level segmentation of actors and their actions in video content. Different from existing works, which all learn to segment from a fixed vocabulary of actor and action pairs, we infer the segmentation from a natural language input sentence. This allows to distinguish betw…

Cited by 197SourcePDFScholar
2018

Real-World Repetition Estimation by Div, Grad and Curl

CVPR 2018poster

We consider the problem of estimating repetition in video, such as performing push-ups, cutting a melon or playing violin. Existing work shows good results under the assumption of static and stationary periodicity. As realistic video is rarely perfectly static and stationary, the often preferred Fou…

Cited by 77SourcePDFScholar
2017

Tracking by Natural Language Specification

CVPR 2017poster

This paper strives to track a target object in a video. Rather than specifying the target in the first frame of a video by a bounding box, we propose to track the object based on a natural language specification of the target, which provides a more natural human-machine interaction as well as a mean…

Cited by 206PDFScholar
2015

Active Transfer Learning With Zero-Shot Priors: Reusing Past Datasets for Future Tasks

ICCV 2015poster

How can we reuse existing knowledge, in the form of available datasets, when solving a new and apparently unrelated target task from a set of unlabeled data? In this work we make a first contribution to answer this question in the context of image classification. We frame this quest as an active…

Cited by 88PDFScholar
2015

Objects2action: Classifying and Localizing Actions Without Any Video Example

ICCV 2015poster

The goal of this paper is to recognize actions in video without the need for examples. Different from traditional zero-shot approaches we do not demand the design and specification of attribute classifiers and class-to-attribute mappings to allow for transfer from seen classes to unseen classes. Our…

Cited by 189PDFScholar
2015

What do 15,000 Object Categories Tell Us About Classifying and Localizing Actions?

CVPR 2015poster

This paper contributes to automatic classification and localization of human actions in video. Whereas motion is the key ingredient in modern approaches, we assess the benefits of having objects in the video representation. Rather than considering a handful of carefully selected and localized object…

Cited by 225SourcePDFScholar