← Search

Bernhard Kainz

14 accepted papers

2026

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits

ICML 2026poster

This paper presents STARCaster, an identity-aware spatio-temporal video diffusion model that addresses both speech-driven portrait animation and free-viewpoint talking portrait synthesis, given an identity embedding or reference image, within a unified framework. Existing 2D speech-to-video diffusio…

Cited by 0SourceScholar
2025

Better Tokens for Better 3D: Advancing Vision-Language Modeling in 3D Medical Imaging

NeurIPS 2025poster

Recent progress in vision-language modeling for 3D medical imaging has been fueled by large-scale computed tomography (CT) corpora with paired free-text reports, stronger architectures, and powerful pretrained models. This has enabled applications such as automated report generation and text-conditi…

Cited by 0SourcecodeScholar
2025

From Self-Check to Consensus: Bayesian Strategic Decoding in Large Language Models

NeurIPS 2025poster

Large Language Models exhibit logical inconsistency across multi-turn inference processes, undermining correctness in complex inferential tasks. Challenges arise from ensuring that outputs align with both factual correctness and human intent. Approaches like single-agent reflection and multi-agent d…

Cited by 0SourceScholar
2025

Image Generation Diversity Issues and How to Tame Them

CVPR 2025poster

Generative methods have reached a level of quality that is almost indistinguishable from real data. However, while individual samples may appear unique, generative models often exhibit limitations in covering the full data distribution. Unlike quality issues, diversity problems within generative mod…

2025

NOVA: A Benchmark for Rare Anomaly Localization and Clinical Reasoning in Brain MRI

NeurIPS 2025oral

In many real-world applications, deployed models encounter inputs that differ from the data seen during training. Open-world recognition ensures that such systems remain robust as ever-emerging, previously _unknown_ categories appear and must be addressed without retraining. Foundation and vision-la…

Cited by 0SourceScholar
2025

SpinMeRound: Consistent Multi-View Identity Generation Using Diffusion Models

ICCV 2025poster

Despite recent progress in diffusion models, generating realistic head portraits from novel viewpoints remains a significant challenge in computer vision. Most current approaches are constrained to limited angular ranges, predominantly focusing on frontal or near-frontal views. Moreover, although th…

Cited by 0SourcePDFScholar
2024

Arc2Face: A Foundation Model for ID-Consistent Human Faces

ECCV 2024oral

"This paper presents , an identity-conditioned face foundation model, which, given the ArcFace embedding of a person, can generate diverse photo-realistic images with an unparalleled degree of face similarity than existing models. Despite previous attempts to decode face recognition features into de…

2024

Stability and Generalizability in SDE Diffusion Models with Measure-Preserving Dynamics

NeurIPS 2024poster

Inverse problems describe the process of estimating the causal factors from a set of measurements or data. Mapping of often incomplete or degraded data to parameters is ill-posed, thus data-driven iterative solutions are required, for example when reconstructing clean images from poor signals. Dif…

Cited by 1SourcePDFScholar
2024

Style-Extracting Diffusion Models for Semi-Supervised Histopathology Segmentation

ECCV 2024poster

"Deep learning-based image generation has seen significant advancements with diffusion models, notably improving the quality of generated images. Despite these developments, generating images with unseen characteristics beneficial for downstream tasks has received limited attention. To bridge this g…

2024

Trade-Offs in Fine-Tuned Diffusion Models between Accuracy and Interpretability

AAAI 2024technical

Recent advancements in diffusion models have significantly impacted the trajectory of generative machine learning re-search, with many adopting the strategy of fine-tuning pre-trained models using domain-specific text-to-image datasets. Notably, this method has been readily employed for medical appl…

2023

Foreground-Background Separation through Concept Distillation from Generative Image Foundation Models

ICCV 2023poster

Curating datasets for object segmentation is a difficult task. With the advent of large-scale pre-trained generative models, conditional image generation has been given a significant boost in result quality and ease of use. In this paper, we present a novel method that enables the generation of gene…

Cited by 9PDFcodeScholar
2022

A Variational Bayesian Method for Similarity Learning in Non-Rigid Image Registration

CVPR 2022poster

We propose a novel variational Bayesian formulation for diffeomorphic non-rigid registration of medical images, which learns in an unsupervised way a data-specific similarity metric. The proposed framework is general and may be used together with many existing image registration models. We evaluate…

Cited by 12PDFcodeScholar
2022

Natural Synthetic Anomalies for Self-Supervised Anomaly Detection and Localization

ECCV 2022poster

"We introduce a simple and intuitive self-supervision task, Natural Synthetic Anomalies (NSA), for training an end-to-end model for anomaly detection and localization using only normal training data. NSA integrates Poisson image editing to seamlessly blend scaled patches of various sizes from separa…

2021

Unsupervised Human Pose Estimation Through Transforming Shape Templates

CVPR 2021poster

Human pose estimation is a major computer vision problem with applications ranging from augmented reality and video capture to surveillance and movement tracking. In the medical context, the latter may be an important biomarker for neurological impairments in infants. Whilst many methods exist, thei…

Cited by 53PDFcodeScholar