← Search

Björn Ommer

35 accepted papers

2026

Adapting Self-Supervised Representations as a Latent Space for Efficient Generation

ICLR 2026poster

We introduce Representation Tokenizer (RepTok), a generative modeling framework that represents an image using a single continuous latent token obtained from self-supervised vision transformers. Building on a pre-trained SSL encoder, we fine-tune only the semantic token embedding and pair it with a…

Cited by 0SourcecodeScholar
2026

Denoising, Fast and Slow: Difficulty-Aware Adaptive Sampling for Image Generation

CVPR 2026

Diffusion- and flow-based models usually allocate compute uniformly across space, updating all patches with the same timestep and number of function evaluations. While convenient, this ignores the heterogeneity of natural images: some regions are easy to denoise, whereas others benefit from more ref

Cited by 0SourcecodeScholar
2026

Envisioning the Future, One Step at a Time

CVPR 2026

Accurately anticipating how complex, diverse scenes will evolve requires models that represent uncertainty, simulate along extended interaction chains, and efficiently explore many plausible futures. Yet most existing approaches rely on dense video or latent-space prediction, expending substantial c

Cited by 0SourcecodeScholar
2026

Guiding Token-Sparse Diffusion Models

CVPR 2026

Diffusion models deliver high quality in image synthesis but remain expensive during training and inference. Recent works have leveraged the inherent redundancy in visual content to make training more affordable by training only on a subset of visual information. While these methods were successful

Cited by 0SourcecodeScholar
2026

Learning Long-term Motion Embeddings for Efficient Kinematics Generation

CVPR 2026

Understanding and predicting motion is a fundamental component of visual intelligence. Although modern video models exhibit strong comprehension of scene dynamics, exploring multiple possible futures through full video synthesis remains prohibitively inefficient. We model scene dynamics orders of ma

Cited by 0SourcecodeScholar
2026

Probabilistic Precipitation Nowcasting with Rectified Flow Transformers

CVPR 2026

Accurate weather forecasts are essential across various domains and are safety-critical in extreme weather conditions. Compared to simulation-based forecasting, data-driven approaches show greater efficiency, enabling short-term, high-resolution nowcasting. In particular, diffusion models proved eff

Cited by 0SourcecodeScholar
2026

Purrception: Variational Flow Matching for Vector-Quantized Image Generation

ICLR 2026poster

We introduce Purrception, a variational flow matching approach for vector-quantized image generation that provides explicit categorical supervision while maintaining continuous transport dynamics. Our method adapts Variational Flow Matching to vector-quantized latents by learning categorical posteri…

Cited by 0SourceScholar
2025

CAGE: Unsupervised Visual Composition and Animation for Controllable Video Generation

AAAI 2025technical

The field of video generation has expanded significantly in recent years, with controllable and compositional video generation garnering considerable interest. Most methods rely on leveraging annotations such as text, objects' bounding boxes, and motion cues, which require substantial human effort a…

2025

CleanDIFT: Diffusion Features without Noise

CVPR 2025poster

Internal features from large-scale pre-trained diffusion models have recently been established as powerful semantic descriptors for a wide range of downstream tasks. Works that use these features generally need to add noise to images before passing them through the model to obtain the semantic featu…

2025

Continuous, Subject-Specific Attribute Control in T2I Models by Identifying Semantic Directions

CVPR 2025poster

Recent advances in text-to-image (T2I) diffusion models have significantly improved the quality of generated images. However, providing efficient control over individual subjects, particularly the attributes characterizing them, remains a key challenge. While existing methods have introduced mechani…

2025

DepthFM: Fast Generative Monocular Depth Estimation with Flow Matching

AAAI 2025technical

Current discriminative depth estimation methods often produce blurry artifacts, while generative approaches suffer from slow sampling due to curvatures in the noise-to-depth transport. Our method addresses these challenges by framing depth estimation as a direct transport between image and depth dis…

2025

Diff2Flow: Training Flow Matching Models via Diffusion Model Alignment

CVPR 2025poster

Diffusion models have revolutionized generative tasks through high-fidelity outputs, yet flow matching (FM) offers faster inference and empirical performance gains. However, current foundation FM models are computationally prohibitive for finetuning, while diffusion models like Stable Diffusion bene…

2025

DisMo: Disentangled Motion Representations for Open-World Motion Transfer

NeurIPS 2025spotlight

Recent advances in text-to-video (T2V) and image-to-video (I2V) models, have enabled the creation of visually compelling and dynamic videos from simple textual descriptions or initial frames. However, these models often fail to provide an explicit representation of motion separate from content, limi…

Cited by 0SourcecodeScholar
2025

Does VLM Classification Benefit from LLM Description Semantics?

AAAI 2025technical

Accurately describing images with text is a foundation of explainable AI. Vision-Language Models (VLMs) like CLIP have recently addressed this by aligning images and texts in a shared embedding space, expressing semantic similarities between vision and language embeddings. VLM classification can be…

2025

Latent Drifting in Diffusion Models for Counterfactual Medical Image Synthesis

CVPR 2025highlight

Scaling by training on large datasets has been shown to enhance the quality and fidelity of image generation and manipulation with diffusion models; however, such large datasets are not always accessible in medical imaging due to cost and privacy issues, which contradicts one of the main application…

Cited by 0SourcePDFScholar
2025

SCFlow: Implicitly Learning Style and Content Disentanglement with Flow Models

ICCV 2025poster

Explicitly disentangling style and content in vision models remains challenging due to their semantic overlap and the subjectivity of human perception. Existing methods propose separation through generative or discriminative objectives, but they still face the inherent ambiguity of disentangling int…

2025

Stochastic Interpolants for Revealing Stylistic Flows across the History of Art

ICCV 2025accepted

Generative models have made rapid progress in content creation, particularly in synthesizing artworks and capturing stylistic variation. However, most methods operate at the level of individual images, limiting their ability to reveal broader stylistic trends and temporal transitions. We address thi…

2025

TREAD: Token Routing for Efficient Architecture-agnostic Diffusion Training

ICCV 2025poster

Diffusion models have emerged as the mainstream approach for visual generation. However, these models typically suffer from sample inefficiency and high training costs. Consequently, methods for efficient finetuning, inference and personalization were quickly adopted by the community. However, train…

2025

What If: Understanding Motion Through Sparse Interactions

ICCV 2025poster

Understanding the dynamics of a physical scene involves reasoning about the diverse ways it can potentially change, especially as a result of local interactions. We present the Flow Poke Transformer (FPT), a novel framework for directly predicting the distribution of local motion, conditioned on spa…

Cited by 0SourcePDFScholar
2024

FMBoost: Boosting Latent Diffusion with Flow Matching

ECCV 2024oral

"Visual synthesis has recently seen significant leaps in performance, largely due to breakthroughs in generative models. Diffusion models have been a key enabler, as they excel in image diversity. However, this comes at the cost of slow training and synthesis, which is only partially alleviated by l…

Cited by 0SourcePDFScholar
2024

Stable-Pose: Leveraging Transformers for Pose-Guided Text-to-Image Generation

NeurIPS 2024poster

Controllable text-to-image (T2I) diffusion models have shown impressive performance in generating high-quality visual content through the incorporation of various conditions. Current methods, however, exhibit limited performance when guided by skeleton human poses, especially in complex pose conditi…

2023

Cross-Image-Attention for Conditional Embeddings in Deep Metric Learning

CVPR 2023poster

Learning compact image embeddings that yield semantic similarities between images and that generalize to unseen test classes, is at the core of deep metric learning (DML). Finding a mapping from a rich, localized image feature map onto a compact embedding vector is challenging: Although similarity e…

Cited by 8SourcePDFScholar
2022

High-Resolution Image Synthesis With Latent Diffusion Models

CVPR 2022oral

By decomposing the image formation process into a sequential application of denoising autoencoders, diffusion models (DMs) achieve state-of-the-art synthesis results on image data and beyond. Additionally, their formulation allows for a guiding mechanism to control the image generation process witho…

Cited by 18394PDFcodeScholar
2022

Retrieval-Augmented Diffusion Models

NeurIPS 2022accept

Novel architectures have recently improved generative image synthesis leading to excellent visual quality in various tasks. Much of this success is due to the scalability of these architectures and hence caused by a dramatic increase in model complexity and in the computational resources invested in…

Cited by 162SourcePDFScholar
2021

Characterizing Generalization under Out-Of-Distribution Shifts in Deep Metric Learning

NeurIPS 2021poster

Deep Metric Learning (DML) aims to find representations suitable for zero-shot transfer to a priori unknown test distributions. However, common evaluation protocols only test a single, fixed data split in which train and test classes are assigned randomly. More realistic evaluations should consider…

Cited by 27SourcePDFScholar
2021

ImageBART: Bidirectional Context with Multinomial Diffusion for Autoregressive Image Synthesis

NeurIPS 2021poster

Autoregressive models and their sequential factorization of the data likelihood have recently demonstrated great potential for image representation and synthesis. Nevertheless, they incorporate image context in a linear 1D order by attending only to previously synthesized image patches above or to t…

Cited by 171SourcePDFScholar
2021

SLIM: Self-Supervised LiDAR Scene Flow and Motion Segmentation

ICCV 2021poster

Recently, several frameworks for self-supervised learning of 3D scene flow on point clouds have emerged. Scene flow inherently separates every scene into multiple moving agents and a large class of points following a single rigid sensor motion. However, existing methods do not leverage this property…

Cited by 113PDFScholar
2021

Shape or Texture: Understanding Discriminative Features in CNNs

ICLR 2021poster

Contrasting the previous evidence that neurons in the later layers of a Convolutional Neural Network (CNN) respond to complex object shapes, recent studies have shown that CNNs actually exhibit a 'texture bias': given an image with both texture and shape cues (e.g., a stylized image), a CNN is biase…

Cited by 89SourcePDFScholar
2021

VelocityNet: Motion-Driven Feature Aggregation for 3D Object Detection in Point Cloud Sequences

ICRA 2021poster

The most successful methods for LiDAR-based 3D object detection use sequences of point clouds in order to exploit the increased data density through temporal aggregation. However, common aggregation methods are rarely able to capture fast-moving objects appropriately. These objects are displaced by…

Cited by 5SourceScholar
2021

iPOKE: Poking a Still Image for Controlled Stochastic Video Synthesis

ICCV 2021poster

How would a static scene react to a local poke? What are the effects on other parts of an object if you could locally push it? There will be distinctive movement, despite evident variations caused by the stochastic nature of our world. These outcomes are governed by the characteristic kinematics of…

Cited by 39PDFScholar
2020

DiVA: Diverse Visual Feature Aggregation for Deep Metric Learning

ECCV 2020poster

Visual Similarity plays an important role in many computer vision applications. Deep metric learning (DML) is a powerful framework for learning such similarities which not only generalize from training data to identically distributed test distributions, but in particular also translate to unknown te…

2020

Making Sense of CNNs: Interpreting Deep Representations & Their Invariances with INNs

ECCV 2020poster

To tackle increasingly complex tasks, it has become an essential ability of neural networks to learn abstract representations. These task-specific representations and, particularly, the invariances they capture turn neural networks into black box models that lack interpretability. To open such a bla…

2018

A Variational U-Net for Conditional Appearance and Shape Generation

CVPR 2018poster

Deep generative models have demonstrated great performance in image synthesis. However, results deteriorate in case of spatial deformations, since they generate images of objects directly, rather than modeling the intricate interplay of their inherent shape and appearance. We present a conditional U…