← Search

Jan-jakob Sonke

14 accepted papers

2026

Distributional Vision-Language Alignment by Cauchy-Schwarz Divergence

ICLR 2026poster

Vision-language alignment is crucial for various downstream tasks such as cross-modal generation and retrieval. Previous multimodal approaches like CLIP utilize InfoNCE to maximize mutual information, primarily aligning pairwise samples across modalities while overlooking distributional differences.…

Cited by 0SourceScholar
2026

Towards Uniformity and Alignment for Multimodal Representation Learning

ICML 2026poster

Multimodal representation learning aims to construct a shared embedding space in which heterogeneous modalities are semantically aligned. Despite strong empirical results, InfoNCE-based objectives introduce inherent conflicts that yield distribution gaps across modalities. In this work, we identify …

Cited by 0SourceScholar
2025

CaPo: Cooperative Plan Optimization for Efficient Embodied Multi-Agent Cooperation

ICLR 2025poster

In this work, we address the cooperation problem among large language model (LLM) based embodied agents, where agents must cooperate to achieve a common goal. Previous methods often execute actions extemporaneously and incoherently, without long-term strategic and cooperative planning, leading to r…

2025

Probabilistic Interactive 3D Segmentation with Hierarchical Neural Processes

ICML 2025poster

Interactive 3D segmentation has emerged as a promising solution for generating accurate object masks in complex 3D scenes by incorporating user-provided clicks. However, two critical challenges remain underexplored: (1) effectively generalizing from sparse user clicks to produce accurate segmentatio…

Cited by 0SourcePDFScholar
2025

Probabilistic Prototype Calibration of Vision-language Models for Generalized Few-shot Semantic Segmentation

ICCV 2025poster

Generalized Few-Shot Semantic Segmentation (GFSS) aims to extend a segmentation model to novel classes with only a few annotated examples while maintaining performance on base classes. Recently, pretrained vision-language models (VLMs) such as CLIP have been leveraged in GFSS to improve generalizati…

2024

Click Prompt Learning with Optimal Transport for Interactive Segmentation

ECCV 2024poster

"Click-based interactive segmentation aims to segment target objects conditioned on user-provided clicks. Existing methods typically interpret user intention by learning multiple click prompts to generate corresponding prompt-activated masks, and selecting one from these masks. However, directly mat…

2024

Domain Adaptation with Cauchy-Schwarz Divergence

UAI 2024poster

Domain adaptation aims to use training data from one or multiple source domains to learn a hypothesis that can be generalized to a different, but related, target domain. As such, having a reliable measure for evaluating the discrepancy of both marginal and conditional distributions is crucial. We in…

2024

How to Train Neural Field Representations: A Comprehensive Study and Benchmark

CVPR 2024poster

Neural fields (NeFs) have recently emerged as a versatile method for modeling signals of various modalities including images shapes and scenes. Subsequently a number of works have explored the use of NeFs as representations for downstream tasks e.g. classifying an image based on the parameters of a…

2024

Kandinsky Conformal Prediction: Efficient Calibration of Image Segmentation Algorithms

CVPR 2024poster

Image segmentation algorithms can be understood as a collection of pixel classifiers for which the outcomes of nearby pixels are correlated. Classifier models can be calibrated using Inductive Conformal Prediction but this requires holding back a sufficiently large calibration dataset for computing…

2024

Space-Time Continuous PDE Forecasting using Equivariant Neural Fields

NeurIPS 2024poster

Recently, Conditional Neural Fields (NeFs) have emerged as a powerful modelling paradigm for PDEs, by learning solutions as flows in the latent space of the Conditional NeF. Although benefiting from favourable properties of NeFs such as grid-agnosticity and space-time-continuous dynamics modelling,…

Cited by 7SourcePDFScholar
2024

Task-Driven Wavelets using Constrained Empirical Risk Minimization

CVPR 2024poster

Deep Neural Networks (DNNs) are widely used for their ability to effectively approximate large classes of functions. This flexibility however makes the strict enforcement of constraints on DNNs a difficult problem. In contexts where it is critical to limit the function space to which certain network…

Cited by 0SourcePDFScholar
2023

Modelling Long Range Dependencies in $N$D: From Task-Specific to a General Purpose CNN

ICLR 2023poster

Performant Convolutional Neural Network (CNN) architectures must be tailored to specific tasks in order to consider the length, resolution, and dimensionality of the input data. In this work, we tackle the need for problem-specific CNN architectures. We present the Continuous Convolutional Neural Ne…

2022

Dynamic Prototype Convolution Network for Few-Shot Semantic Segmentation

CVPR 2022poster

The key challenge for few-shot semantic segmentation (FSS) is how to tailor a desirable interaction among support and query features and/or their prototypes, under the episodic training scenario. Most existing FSS methods implement such support/query interactions by solely leveraging \it plain oper…

Cited by 111PDFScholar
2022

Recurrent Variational Network: A Deep Learning Inverse Problem Solver Applied to the Task of Accelerated MRI Reconstruction

CVPR 2022poster

Magnetic Resonance Imaging can produce detailed images of the anatomy and physiology of the human body that can assist doctors in diagnosing and treating pathologies such as tumours. However, MRI suffers from very long acquisition times that make it susceptible to patient motion artifacts and limit…

Cited by 82PDFcodeScholar