← Search

Stephan Günnemann

111 accepted papers

2026

A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness

ICML 2026poster

Automated \enquote{LLM-as-a-Judge} frameworks have become the de facto standard for scalable evaluation across natural language processing. For instance, in safety evaluation, these judges are relied upon to evaluate harmfulness in order to benchmark the robustness of safety against adversarial atta…

Cited by 0SourceScholar
2026

Certifying Graph Neural Networks Against Label and Structure Poisoning

ICML 2026poster

Robust machine learning for graph-structured data has made significant progress against test-time attacks, yet certified robustness to poisoning – where adversaries manipulate the training data – remains largely underexplored. For image data, state-of-the-art poisoning certificates rely on partition…

Cited by 0SourceScholar
2026

Derivative Informed Learning of Exchange-Correlation Functionals

ICML 2026poster

Machine-learned (ML) XC-functionals promise improved accuracy, but overfit to training energies and basis sets without proper regularization. We introduce Derivative Informed XC-Loss (DI-Loss), a loss that regularizes ML-XC training by supervising energy gradients on the Grassmannian of density matr…

Cited by 0SourceScholar
2026

Discrete Bayesian Sample Inference for Graph Generation

ICLR 2026poster

Generating graph-structured data is crucial in applications such as molecular generation, knowledge graphs, and network analysis. However, their discrete, unordered nature makes them difficult for traditional generative models, leading to the rise of discrete diffusion and flow matching models. In t…

Cited by 0SourcecodeScholar
2026

Excited Pfaffians: Generalized Neural Wave Functions Across Structure and State

ICML 2026spotlight

Neural-network wave functions in Variational Monte Carlo (VMC) have achieved great success in accurately representing both ground and excited states. However, achieving sufficient numerical accuracy of state overlaps requires growing the number of Monte Carlo samples, and consequently computational …

Cited by 0SourceScholar
2026

Flow-Based Density Ratio Estimation for Intractable Distributions with Applications in Genomics

ICML 2026poster

Estimating density ratios between pairs of intractable data distributions is a core problem in probabilistic modeling, enabling principled comparisons of sample likelihoods under different data-generating processes across conditions and covariates. While exact-likelihood models such as normalizing f…

Cited by 0SourceScholar
2026

Model Collapse Is Not a Bug but a Feature in Machine Unlearning for LLMs

ICLR 2026poster

Current unlearning methods for LLMs optimize on the private information they seek to remove by incorporating it into their fine-tuning data. We argue this not only risks reinforcing exposure to sensitive data, it also fundamentally contradicts the principle of minimizing its use. As a remedy, we pro…

Cited by 0SourcecodeScholar
2026

Position: LLM-Safety Evaluations Lack Robustness

ICML 2026poster

In this position paper, we argue that current safety alignment research efforts for large language models are hindered by many intertwined sources of noise, such as small datasets, methodological inconsistencies, and unreliable evaluation setups. This can, at times, make it impossible to evaluate an…

Cited by 0SourceScholar
2026

Sampling-aware Adversarial Attacks Against Large Language Models

ICLR 2026poster

To guarantee safe and robust deployment of large language models (LLMs) at scale, it is critical to accurately assess their adversarial robustness. Existing adversarial attacks typically target harmful responses in single-point greedy generations, overlooking the inherently stochastic nature of LLMs…

Cited by 0SourceScholar
2025

A Probabilistic Perspective on Unlearning and Alignment for Large Language Models

ICLR 2025oral

Comprehensive evaluation of Large Language Models (LLMs) is an open research problem. Existing evaluations rely on deterministic point estimates generated via greedy decoding. However, we find that deterministic evaluations fail to capture the whole output distribution of a model, yielding inaccurat…

2025

Consistent Sampling and Simulation: Molecular Dynamics with Energy-Based Diffusion Models

NeurIPS 2025poster

In recent years, diffusion models trained on equilibrium molecular distributions have proven effective for sampling biomolecules. Beyond direct sampling, the score of such a model can also be used to derive the forces that act on molecular systems. However, while classical diffusion sampling usually…

Cited by 0SourcecodeScholar
2025

Efficient Time Series Processing for Transformers and State-Space Models through Token Merging

ICML 2025poster

Despite recent advances in subquadratic attention mechanisms or state-space models, processing long token sequences still imposes significant computational requirements. Token merging has emerged as a solution to increase computational efficiency in computer vision architectures. In this work, we pe…

Cited by 5SourcePDFScholar
2025

Enforcing Latent Euclidean Geometry in Single-Cell VAEs for Manifold Interpolation

ICML 2025spotlight

Latent space interpolations are a powerful tool for navigating deep generative models in applied settings. An example is single-cell RNA sequencing, where existing methods model cellular state transitions as latent space interpolations with variational autoencoders, often assuming linear shifts and…

Cited by 0SourcePDFScholar
2025

Exact Certification of (Graph) Neural Networks Against Label Poisoning

ICLR 2025spotlight

Machine learning models are highly vulnerable to label flipping, i.e., the adversarial modification (poisoning) of training labels to compromise performance. Thus, deriving robustness certificates is important to guarantee that test predictions remain unaffected and to understand worst-case robustne…

2025

Flow Matching with Gaussian Process Priors for Probabilistic Time Series Forecasting

ICLR 2025poster

Recent advancements in generative modeling, particularly diffusion models, have opened new directions for time series modeling, achieving state-of-the-art performance in forecasting and synthesis. However, the reliance of diffusion-based models on a simple, fixed prior complicates the generative pro…

Cited by 2SourcePDFScholar
2025

GeoDiffusion: A Training-Free Framework for Accurate 3D Geometric Conditioning in Image Generation

ICCV 2025poster

Precise geometric control in image generation is essential for fields like engineering & product design and creative industries to control 3D object features accurately in 2D image space. Traditional 3D editing approaches are time-consuming and demand specialized skills, while current image-based ge…

Cited by 0SourcePDFScholar
2025

Graph Neural Networks for Edge Signals: Orientation Equivariance and Invariance

ICLR 2025poster

Many applications in traffic, civil engineering, or electrical engineering revolve around edge-level signals. Such signals can be categorized as inherently directed, for example, the water flow in a pipe network, and undirected, like the diameter of a pipe. Topological methods model edge signals wit…

Cited by 1SourcePDFScholar
2025

Joint Out-of-Distribution Filtering and Data Discovery Active Learning

CVPR 2025poster

As the data demand for deep learning models increases, active learning (AL) becomes essential to strategically select samples for labeling, which maximizes data efficiency and reduces training costs. Real-world scenarios necessitate the consideration of incomplete data knowledge within AL. Prior wor…

Cited by 1SourcePDFScholar
2025

Joint Relational Database Generation via Graph-Conditional Diffusion Models

NeurIPS 2025poster

Building generative models for relational databases (RDBs) is important for many applications, such as privacy-preserving data release and augmenting real datasets. However, most prior works either focus on single-table generation or adapt single-table models to the multi-table setting by relying on…

Cited by 0SourcecodeScholar
2025

Lift Your Molecules: Molecular Graph Generation in Latent Euclidean Space

ICLR 2025poster

We introduce a new framework for 2D molecular graph generation using 3D molecule generative models. Our Synthetic Coordinate Embedding (SyCo) framework maps 2D molecular graphs to 3D Euclidean point clouds via synthetic coordinates and learns the inverse map using an E($n$)-Equivariant Graph Neural…

Cited by 1SourcePDFScholar
2025

MAGNet: Motif-Agnostic Generation of Molecules from Scaffolds

ICLR 2025spotlight

Recent advances in machine learning for molecules exhibit great potential for facilitating drug discovery from in silico predictions. Most models for molecule generation rely on the decomposition of molecules into frequently occurring substructures (motifs), from which they generate novel compounds.…

Cited by 0SourcePDFScholar
2025

Modeling Microenvironment Trajectories on Spatial Transcriptomics with NicheFlow

NeurIPS 2025poster

Understanding the evolution of cellular microenvironments in spatiotemporal data is essential for deciphering tissue development and disease progression. While experimental techniques like spatial transcriptomics now enable high-resolution mapping of tissue organization across space and time, curren…

Cited by 0SourceScholar
2025

Prior2Former - Evidential Modeling of Mask Transformers for Assumption-Free Open-World Panoptic Segmentation

ICCV 2025poster

In panoptic segmentation, individual instances must be separated within semantic classes. As state-of-the-art methods rely on a pre-defined set of classes, they struggle with novel categories and out-of-distribution (OOD) data. This is particularly problematic in safety-critical applications, such a…

Cited by 0SourcePDFScholar
2025

Privacy Amplification by Structured Subsampling for Deep Differentially Private Time Series Forecasting

ICML 2025spotlight

Many forms of sensitive data, such as web traffic, mobility data, or hospital occupancy, are inherently sequential. The standard method for training machine learning models while ensuring privacy for units of sensitive information, such as individual hospital visits, is differentially private stocha…

Cited by 0SourcePDFScholar
2025

Provably Reliable Conformal Prediction Sets in the Presence of Data Poisoning

ICLR 2025spotlight

Conformal prediction provides model-agnostic and distribution-free uncertainty quantification through prediction sets that are guaranteed to include the ground truth with any user-specified probability. Yet, conformal prediction is not reliable under poisoning attacks where adversaries manipulate bo…

Cited by 0SourcePDFScholar
2025

REINFORCE Adversarial Attacks on Large Language Models: An Adaptive, Distributional, and Semantic Objective

ICML 2025poster

To circumvent the alignment of large language models (LLMs), current optimization-based adversarial attacks usually craft adversarial prompts by maximizing the likelihood of a so-called affirmative response. An affirmative response is a manually designed start of a harmful answer to an inappropriate…

2025

The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence

ICML 2025poster

The safety alignment of large language models (LLMs) can be circumvented through adversarially crafted inputs, yet the mechanisms by which these attacks bypass safety barriers remain poorly understood. Prior work suggests that a *single* refusal direction in the model's activation space determines w…

Cited by 0SourcePDFScholar
2025

TreeGen: A Bayesian Generative Model for Hierarchies

NeurIPS 2025poster

In this work, we introduce TreeGen, a novel generative framework modeling distributions over hierarchies. We extend Bayesian Flow Networks (BFNs) to enable transitions between probabilistic and discrete hierarchies parametrized via categorical distributions. Our proposed scheduler provides smooth an…

Cited by 0SourceScholar
2025

UnHiPPO: Uncertainty-aware Initialization for State Space Models

ICML 2025poster

State space models are emerging as a dominant model class for sequence problems with many relying on the HiPPO framework to initialize their dynamics. However, HiPPO fundamentally assumes data to be noise-free; an assumption often violated in practice. We extend the HiPPO theory with measurement noi…

Cited by 0SourcePDFScholar
2025

Uncertainty Estimation for Heterophilic Graphs Through the Lens of Information Theory

ICML 2025poster

While uncertainty estimation for graphs recently gained traction, most methods rely on homophily and deteriorate in heterophilic settings. We address this by analyzing message passing neural networks from an information-theoretic perspective and developing a suitable analog to data processing in…

Cited by 0SourcePDFScholar
2025

Unlocking Point Processes through Point Set Diffusion

ICLR 2025poster

Point processes model the distribution of random point sets in mathematical spaces, such as spatial and temporal domains, with applications in fields like seismology, neuroscience, and economics. Existing statistical and machine learning models for point processes are predominantly constrained by th…

Cited by 1SourcePDFScholar
2025

What Expressivity Theory Misses: Message Passing Complexity for GNNs

NeurIPS 2025spotlight

Expressivity theory, characterizing which graphs a GNN can distinguish, has become the predominant framework for analyzing GNNs, with new models striving for higher expressivity. However, we argue that this focus is misguided: First, higher expressivity is not necessary for most real-world tasks as…

Cited by 0SourceScholar
2024

Deep Sensor Fusion with Constraint Safety Bounds for High Precision Localization

IROS 2024poster

In mobile robotics, particularly in autonomous driving, localization is one of the key challenges for navigation and planning. For safe operation in the open world where vulnerable participants are present, precise and guaranteed safe localization is required. While current classical fusion approach…

Cited by 2SourceScholar
2024

Efficient Adversarial Training in LLMs with Continuous Attacks

NeurIPS 2024spotlight

Large language models (LLMs) are vulnerable to adversarial attacks that can bypass their safety guardrails. In many domains, adversarial training has proven to be one of the most promising methods to reliably improve robustness against such attacks. Yet, in the context of LLMs, current methods for a…

2024

Energy-based Epistemic Uncertainty for Graph Neural Networks

NeurIPS 2024spotlight

In domains with interdependent data, such as graphs, quantifying the epistemic uncertainty of a Graph Neural Network (GNN) is challenging as uncertainty can arise at different structural scales. Existing techniques neglect this issue or only distinguish between structure-aware and structure-agnostic…

Cited by 1SourcePDFScholar
2024

Expressivity and Generalization: Fragment-Biases for Molecular GNNs

ICML 2024oral

Although recent advances in higher-order Graph Neural Networks (GNNs) improve the theoretical expressiveness and molecular property predictive performance, they often fall short of the empirical performance of models that explicitly use fragment information as inductive bias. However, for these appr…

Cited by 5SourcePDFScholar
2024

From Zero to Turbulence: Generative Modeling for 3D Flow Simulation

ICLR 2024poster

Simulations of turbulent flows in 3D are one of the most expensive simulations in computational fluid dynamics (CFD). Many works have been written on surrogate models to replace numerical solvers for fluid flows with faster, learned, autoregressive models. However, the intricacies of turbulence in t…

2024

Generalized Synchronized Active Learning for Multi-Agent-Based Data Selection on Mobile Robotic Systems

RA-L 2024

In mobile robotics, perception in uncontrolled environments like autonomous driving is a central hurdle. Existing active learning frameworks can help enhance perception by efficiently selecting data samples for labeling, but they are often constrained by the necessity of full data availability in da

Cited by 6SourceScholar
2024

Guaranteeing Robustness Against Real-World Perturbations In Time Series Classification Using Conformalized Randomized Smoothing

UAI 2024poster

Certifying the robustness of machine learning models against domain shifts and input space perturbations is crucial for many applications, where high risk decisions are based on the model’s predictions. Techniques such as randomized smoothing have partially addressed this issues with a focus on adve…

Cited by 0SourcePDFScholar
2024

Shaving Weights with Occam's Razor: Bayesian Sparsification for Neural Networks using the Marginal Likelihood

NeurIPS 2024poster

Neural network sparsification is a promising avenue to save computational time and memory costs, especially in an age where many successful AI models are becoming too large to naively deploy on consumer hardware. While much work has focused on different weight pruning criteria, the overall sparsifia…

2024

Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

NeurIPS 2024poster

Current research in adversarial robustness of LLMs focuses on \textit{discrete} input manipulations in the natural language space, which can be directly transferred to \textit{closed-source} models. However, this approach neglects the steady progression of \textit{open-source} models. As open-source…

2024

Spatio-Spectral Graph Neural Networks

NeurIPS 2024poster

Spatial Message Passing Graph Neural Networks (MPGNNs) are widely used for learning on graph-structured data. However, key limitations of *ℓ*-step MPGNNs are that their "receptive field" is typically limited to the *ℓ*-hop neighborhood of a node and that information exchange between distant nodes is…

2024

Uncertainty for Active Learning on Graphs

ICML 2024poster

Uncertainty Sampling is an Active Learning strategy that aims to improve the data efficiency of machine learning models by iteratively acquiring labels of data points with the highest uncertainty. While it has proven effective for independent data its applicability to graphs remains under-explored.…

Cited by 10SourcePDFScholar
2024

Unified Guidance for Geometry-Conditioned Molecular Generation

NeurIPS 2024poster

Effectively designing molecular geometries is essential to advancing pharmaceutical innovations, a domain, which has experienced great attention through the success of generative models and, in particular, diffusion models. However, current molecular diffusion models are tailored towards a specific…

Cited by 2SourcePDFScholar
2024

Unified Mechanism-Specific Amplification by Subsampling and Group Privacy Amplification

NeurIPS 2024poster

Amplification by subsampling is one of the main primitives in machine learning with differential privacy (DP): Training a model on random batches instead of complete datasets results in stronger privacy. This is traditionally formalized via mechanism-agnostic subsampling guarantees that express the…

Cited by 2SourcePDFScholar
2023

Add and Thin: Diffusion for Temporal Point Processes

NeurIPS 2023poster

Autoregressive neural networks within the temporal point process (TPP) framework have become the standard for modeling continuous-time event data. Even though these models can expressively capture event sequences in a one-step-ahead fashion, they are inherently limited for long-term forecasting appl…

Cited by 14SourcePDFScholar
2023

Adversarial Training for Graph Neural Networks: Pitfalls, Solutions, and New Directions

NeurIPS 2023poster

Despite its success in the image domain, adversarial training did not (yet) stand out as an effective defense for Graph Neural Networks (GNNs) against graph structure perturbations. In the pursuit of fixing adversarial training (1) we show and overcome fundamental theoretical as well as practical l…

Cited by 34SourcePDFScholar
2023

Ewald-based Long-Range Message Passing for Molecular Graphs

ICML 2023poster

Neural architectures that learn potential energy surfaces from molecular data have undergone fast improvement in recent years. A key driver of this success is the Message Passing Neural Network (MPNN) paradigm. Its favorable scaling with system size partly relies upon a spatial distance limit on mes…

2023

Localized Randomized Smoothing for Collective Robustness Certification

ICLR 2023top-25%

Models for image segmentation, node classification and many other tasks map a single input to multiple labels. By perturbing this single shared input (e.g. the image) an adversary can manipulate several predictions (e.g. misclassify several pixels). Collective robustness certification is the task of…

Cited by 11SourcePDFScholar
2023

Modeling Temporal Data as Continuous Functions with Stochastic Process Diffusion

ICML 2023poster

Temporal data such as time series can be viewed as discretized measurements of the underlying function. To build a generative model for such data we have to model the stochastic process that governs it. We propose a solution by defining the denoising diffusion model in the function space which also…

Cited by 44SourcePDFScholar
2023

Provable Adversarial Robustness for Group Equivariant Tasks: Graphs, Point Clouds, Molecules, and More

NeurIPS 2023poster

A machine learning model is traditionally considered robust if its prediction remains (almost) constant under input perturbations with small norm. However, real-world tasks like molecular property prediction or point cloud segmentation have inherent equivariances, such as rotation or permutation equ…

Cited by 4SourcePDFScholar
2023

Sampling-free Inference for Ab-Initio Potential Energy Surface Networks

ICLR 2023poster

Recently, it has been shown that neural networks not only approximate the ground-state wave functions of a single molecular system well but can also generalize to multiple geometries. While such generalization significantly speeds up training, each energy evaluation still requires Monte Carlo integr…

2023

Topology-Matching Normalizing Flows for Out-of-Distribution Detection in Robot Learning

CoRL 2023poster

To facilitate reliable deployments of autonomous robots in the real world, Out-of-Distribution (OOD) detection capabilities are often required. A powerful approach for OOD detection is based on density estimation with Normalizing Flows (NFs). However, we find that prior work with NFs attempts to mat…

Cited by 6SourceScholar
2023

Transformers Meet Directed Graphs

ICML 2023poster

Transformers were originally proposed as a sequence-to-sequence model for text but have become vital for a wide range of modalities, including images, audio, video, and undirected graphs. However, transformers for directed graphs are a surprisingly underexplored topic, despite their applicability to…

2023

Uncertainty Estimation for Molecules: Desiderata and Methods

ICML 2023poster

Graph Neural Networks (GNNs) are promising surrogates for quantum mechanical calculations as they establish unprecedented low errors on collections of molecular dynamics (MD) trajectories. Thanks to their fast inference times they promise to accelerate computational chemistry applications. Unfortuna…

Cited by 13SourcePDFScholar
2023

Unveiling the sampling density in non-uniform geometric graphs

ICLR 2023poster

A powerful framework for studying graphs is to consider them as geometric graphs: nodes are randomly sampled from an underlying metric space, and any pair of nodes is connected if their distance is less than a specified neighborhood radius. Currently, the literature mostly focuses on uniform samplin…

Cited by 3SourcePDFScholar
2022

3D Infomax improves GNNs for Molecular Property Prediction

ICML 2022spotlight

Molecular property prediction is one of the fastest-growing applications of deep learning with critical real-world impacts. Although the 3D molecular graph structure is necessary for models to achieve strong performance on many tasks, it is infeasible to obtain 3D structures at the scale required by…

2022

Ab-Initio Potential Energy Surfaces by Pairing GNNs with Neural Wave Functions

ICLR 2022spotlight

Solving the Schrödinger equation is key to many quantum mechanical properties. However, an analytical solution is only tractable for single-electron systems. Recently, neural networks succeeded at modelling wave functions of many-electron systems. Together with the variational Monte-Carlo (VMC) fram…

2022

Are Defenses for Graph Neural Networks Robust?

NeurIPS 2022accept

A cursory reading of the literature suggests that we have made a lot of progress in designing effective adversarial defenses for Graph Neural Networks (GNNs). Yet, the standard methodology has a serious flaw – virtually all of the defenses are evaluated against non-adaptive attacks leading to overly…

Cited by 84SourcePDFScholar
2022

End-to-End Learning of Probabilistic Hierarchies on Graphs

ICLR 2022poster

We propose a novel probabilistic model over hierarchies on graphs obtained by continuous relaxation of tree-based hierarchies. We draw connections to Markov chain theory, enabling us to perform hierarchical clustering by efficient end-to-end optimization of relaxed versions of quality metrics such a…

Cited by 3SourcePDFScholar
2022

Generalization of Neural Combinatorial Solvers Through the Lens of Adversarial Robustness

ICLR 2022poster

End-to-end (geometric) deep learning has seen first successes in approximating the solution of combinatorial optimization problems. However, generating data in the realm of NP-hard/-complete tasks brings practical and theoretical challenges, resulting in evaluation protocols that are too optimistic.…

Cited by 51SourcePDFScholar
2022

Intriguing Properties of Input-Dependent Randomized Smoothing

ICML 2022spotlight

Randomized smoothing is currently considered the state-of-the-art method to obtain certifiably robust classifiers. Despite its remarkable performance, the method is associated with various serious problems such as “certified accuracy waterfalls”, certification vs. accuracy trade-off, or even fairnes…

Cited by 31SourcePDFScholar
2022

Learning the Dynamics of Physical Systems from Sparse Observations with Finite Element Networks

ICLR 2022spotlight

We propose a new method for spatio-temporal forecasting on arbitrarily distributed points. Assuming that the observed system follows an unknown partial differential equation, we derive a continuous-time model for the dynamics of the data via the finite element method. The resulting graph neural netw…

2022

Natural Posterior Network: Deep Bayesian Predictive Uncertainty for Exponential Family Distributions

ICLR 2022spotlight

Uncertainty awareness is crucial to develop reliable machine learning models. In this work, we propose the Natural Posterior Network (NatPN) for fast and high-quality uncertainty estimation for any task where the target distribution belongs to the exponential family. Thus, NatPN finds application fo…

Cited by 83SourcePDFScholar
2022

Predicting Cellular Responses to Novel Drug Perturbations at a Single-Cell Resolution

NeurIPS 2022accept

Single-cell transcriptomics enabled the study of cellular heterogeneity in response to perturbations at the resolution of individual cells. However, scaling high-throughput screens (HTSs) to measure cellular responses for many drugs remains a challenge due to technical limitations and, more importan…

2022

Randomized Message-Interception Smoothing: Gray-box Certificates for Graph Neural Networks

NeurIPS 2022accept

Randomized smoothing is one of the most promising frameworks for certifying the adversarial robustness of machine learning models, including Graph Neural Networks (GNNs). Yet, existing randomized smoothing certificates for GNNs are overly pessimistic since they treat the model as a black box, ignori…

Cited by 25SourcePDFScholar
2022

Winning the Lottery Ahead of Time: Efficient Early Network Pruning

ICML 2022spotlight

Pruning, the task of sparsifying deep neural networks, received increasing attention recently. Although state-of-the-art pruning methods extract highly sparse models, they neglect two main challenges: (1) the process of finding these sparse models is often very expensive; (2) unstructured pruning do…

2021

Collective Robustness Certificates: Exploiting Interdependence in Graph Neural Networks

ICLR 2021poster

In tasks like node classification, image segmentation, and named-entity recognition we have a classifier that simultaneously outputs multiple predictions (a vector of labels) based on a single input, i.e. a single graph, image, or document respectively. Existing adversarial robustness certificates c…

Cited by 33SourcePDFScholar
2021

Completing the Picture: Randomized Smoothing Suffers from the Curse of Dimensionality for a Large Family of Distributions

AISTATS 2021poster

Randomized smoothing is currently the most competitive technique for providing provable robustness guarantees. Since this approach is model-agnostic and inherently scalable we can certify arbitrary classifiers. Despite its success, recent works show that for a small class of i.i.d. distributions, th…

2021

Detecting Anomalous Event Sequences with Temporal Point Processes

NeurIPS 2021poster

Automatically detecting anomalies in event data can provide substantial value in domains such as healthcare, DevOps, and information security. In this paper, we frame the problem of detecting anomalous continuous-time event sequences as out-of-distribution (OOD) detection for temporal point processe…

Cited by 18SourcePDFScholar
2021

Directional Message Passing on Molecular Graphs via Synthetic Coordinates

NeurIPS 2021poster

Graph neural networks that leverage coordinates via directional message passing have recently set the state of the art on multiple molecular property prediction tasks. However, they rely on atom position information that is often unavailable, and obtaining it is usually prohibitively expensive or ev…

Cited by 53SourcePDFScholar
2021

Evaluating Robustness of Predictive Uncertainty Estimation: Are Dirichlet-based Models Reliable?

ICML 2021spotlight

Dirichlet-based uncertainty (DBU) models are a recent and promising class of uncertainty-aware models. DBU models predict the parameters of a Dirichlet distribution to provide fast, high-quality uncertainty estimates alongside with class predictions. In this work, we present the first large-scale, i…

2021

GemNet: Universal Directional Graph Neural Networks for Molecules

NeurIPS 2021poster

Effectively predicting molecular interactions has the potential to accelerate molecular dynamics by multiple orders of magnitude and thus revolutionize chemical simulations. Graph neural networks (GNNs) have recently shown great successes for this task, overtaking classical methods based on fixed mo…

2021

Graph Posterior Network: Bayesian Predictive Uncertainty for Node Classification

NeurIPS 2021poster

The interdependence between nodes in graphs is key to improve class prediction on nodes, utilized in approaches like Label Probagation (LP) or in Graph Neural Networks (GNNs). Nonetheless, uncertainty estimation for non-independent node-level predictions is under-explored. In this work, we explore…

2021

Language-Agnostic Representation Learning of Source Code from Structure and Context

ICLR 2021poster

Source code (Context) and its parsed abstract syntax tree (AST; Structure) are two complementary representations of the same computer program. Traditionally, designers of machine learning models have relied predominantly either on Structure or Context. We propose a new model, which jointly learns on…

2021

Neural Flows: Efficient Alternative to Neural ODEs

NeurIPS 2021poster

Neural ordinary differential equations describe how values change in time. This is the reason why they gained importance in modeling sequential data, especially when the observations are made at irregular intervals. In this paper we propose an alternative by directly modeling the solution curves - t…

2021

Neural Temporal Point Processes: A Review

IJCAI 2021poster

Temporal point processes (TPP) are probabilistic generative models for continuous-time event sequences. Neural TPPs combine the fundamental ideas from point process literature with deep learning approaches, thus enabling construction of flexible and efficient models. The topic of neural TPPs has att…

Cited by 120SourcePDFScholar
2021

Robustness of Graph Neural Networks at Scale

NeurIPS 2021poster

Graph Neural Networks (GNNs) are increasingly important given their popularity and the diversity of applications. Yet, existing studies of their vulnerability to adversarial attacks rely on relatively small graphs. We address this gap and study how to attack and defend GNNs at scale. We propose two…

2021

Scalable Optimal Transport in High Dimensions for Graph Distances, Embedding Alignment, and More

ICML 2021spotlight

The current best practice for computing optimal transport (OT) is via entropy regularization and Sinkhorn iterations. This algorithm runs in quadratic time as it requires the full pairwise cost matrix, which is prohibitively expensive for large sets of objects. In this work we propose two effective…

Cited by 15SourcePDFScholar
2021

Whole Brain Vessel Graphs: A Dataset and Benchmark for Graph Learning and Neuroscience

NeurIPS 2021poster

Biological neural networks define the brain function and intelligence of humans and other mammals, and form ultra-large, spatial, structured graphs. Their neuronal organization is closely interconnected with the spatial organization of the brain's microvasculature, which supplies oxygen to the neuro…

Cited by 28SourceScholar
2020

Continual Learning with Bayesian Neural Networks for Non-Stationary Data

ICLR 2020poster

This work addresses continual learning for non-stationary data, using Bayesian neural networks and memory-based online variational Bayes. We represent the posterior approximation of the network weights by a diagonal Gaussian distribution and a complementary memory of raw data. This raw data correspo…

Cited by 101SourceScholar
2020

Deep Rao-Blackwellised Particle Filters for Time Series Forecasting

NeurIPS 2020poster

This work addresses efficient inference and learning in switching Gaussian linear dynamical systems using a Rao-Blackwellised particle filter and a corresponding Monte Carlo objective. To improve the forecasting capabilities, we extend this classical model by conditionally linear state-to-switch dyn…

Cited by 44SourcePDFScholar
2020

Efficient Robustness Certificates for Discrete Data: Sparsity-Aware Randomized Smoothing for Graphs, Images and More

ICML 2020poster

Existing techniques for certifying the robustness of models for discrete data either work only for a small class of models or are general at the expense of efficiency or tightness. Moreover, they do not account for sparsity in the input which, as our findings show, is often essential for obtaining n…

Cited by 109SourcePDFScholar
2020

Fast and Flexible Temporal Point Processes with Triangular Maps

NeurIPS 2020oral

Temporal point process (TPP) models combined with recurrent neural networks provide a powerful framework for modeling continuous-time event data. While such models are flexible, they are inherently sequential and therefore cannot benefit from the parallelism of modern hardware. By exploiting the rec…

2020

Posterior Network: Uncertainty Estimation without OOD Samples via Density-Based Pseudo-Counts

NeurIPS 2020poster

Accurate estimation of aleatoric and epistemic uncertainty is crucial to build safe and reliable systems. Traditional approaches, such as dropout and ensemble methods, estimate uncertainty by sampling probability predictions from different submodels, which leads to slow uncertainty estimation at inf…

2019

Failing Loudly: An Empirical Study of Methods for Detecting Dataset Shift

NeurIPS 2019poster

We might hope that when faced with unexpected inputs, well-designed software systems would fire off warnings. Machine learning (ML) systems, however, which depend strongly on properties of their inputs (e.g. the i.i.d. assumption), tend to fail silently. This paper explores the problem of building M…

2019

Predict then Propagate: Graph Neural Networks meet Personalized PageRank

ICLR 2019poster

Neural message passing algorithms for semi-supervised classification on graphs have recently achieved great success. However, for classifying a node these methods only consider nodes that are a few propagation steps away and the size of this utilized neighborhood is hard to extend. In this paper, we…

2019

Uncertainty on Asynchronous Time Event Prediction

NeurIPS 2019spotlight

Asynchronous event sequences are the basis of many applications throughout different industries. In this work, we tackle the task of predicting the next event (given a history), and how this prediction changes with the passage of time. Since at some time points (e.g. predictions far into the future)…

2018

Deep Gaussian Embedding of Graphs: Unsupervised Inductive Learning via Ranking

ICLR 2018poster

Methods that learn representations of nodes in a graph play a critical role in network analysis since they enable many downstream learning tasks. We propose Graph2Gauss - an approach that can efficiently learn versatile node embeddings on large scale (attributed) graphs that show strong performance…

Cited by 885SourcePDFScholar
2018

NetGAN: Generating Graphs via Random Walks

ICML 2018oral

We propose NetGAN - the first implicit generative model for graphs able to mimic real-world networks. We pose the problem of graph generation as learning the distribution of biased random walks over the input graph. The proposed model is based on a stochastic neural network that generates discrete o…

Cited by 505SourcePDFScholar