← Search

Kenji Fukumizu

34 accepted papers

2026

Conditionally Whitened Generative Models for Probabilistic Time Series Forecasting

ICLR 2026poster

Probabilistic forecasting of multivariate time series is challenging due to non-stationarity, inter-variable dependencies, and distribution shifts. While recent diffusion and flow matching models have shown promise, they often ignore informative priors such as conditional means and covariances. In t…

Cited by 0SourcecodeScholar
2025

An Efficient Orlicz-Sobolev Approach for Transporting Unbalanced Measures on a Graph

NeurIPS 2025spotlight

We investigate optimal transport (OT) for measures on graph metric spaces with different total masses. To mitigate the limitations of traditional $L^p$ geometry, Orlicz-Wasserstein (OW) and generalized Sobolev transport (GST) employ \emph{Orlicz geometric structure}, leveraging convex functions to c…

Cited by 0SourceScholar
2025

Compositional simulation-based inference for time series

ICLR 2025poster

Amortized simulation-based inference (SBI) methods train neural networks on simulated data to perform Bayesian inference. While this strategy avoids the need for tractable likelihoods, it often requires a large number of simulations and has been challenging to scale to time series data. Scientific s…

2025

Flow matching achieves almost minimax optimal convergence

ICLR 2025poster

Flow matching (FM) has gained significant attention as a simulation-free generative model. Unlike diffusion models, which are based on stochastic differential equations, FM employs a simpler approach by solving an ordinary differential equation with an initial condition from a normal distribution, t…

Cited by 3SourcePDFScholar
2025

Pairwise Optimal Transports for Training All-to-All Flow-Based Condition Transfer Model

NeurIPS 2025poster

In this paper, we propose a flow-based method for learning all-to-all transfer maps among conditional distributions that approximates pairwise optimal transport. The proposed method addresses the challenge of handling the case of continuous conditions, which often involve a large set of conditions w…

Cited by 0SourceScholar
2024

Neural Fourier Transform: A General Approach to Equivariant Representation Learning

ICLR 2024poster

Symmetry learning has proven to be an effective approach for extracting the hidden structure of data, with the concept of equivariance relation playing the central role. However, most of the current studies are built on architectural theory and corresponding assumptions on the form of data. We pro…

Cited by 5SourcePDFScholar
2023

Controlling Posterior Collapse by an Inverse Lipschitz Constraint on the Decoder Network

ICML 2023poster

Variational autoencoders (VAEs) are one of the deep generative models that have experienced enormous success over the past decades. However, in practice, they suffer from a problem called posterior collapse, which occurs when the posterior distribution coincides, or collapses, with the prior taking…

Cited by 6SourcePDFScholar
2023

Transfer Learning with Affine Model Transformation

NeurIPS 2023poster

Supervised transfer learning has received considerable attention due to its potential to boost the predictive power of machine learning in scenarios where data are scarce. Generally, a given set of source models and a dataset from a target domain are used to adapt the pre-trained models to a target…

2022

$\beta$-Intact-VAE: Identifying and Estimating Causal Effects under Limited Overlap

ICLR 2022poster

As an important problem in causal inference, we discuss the identification and estimation of treatment effects (TEs) under limited overlap; that is, when subjects with certain features belong to a single treatment group. We use a latent variable to model a prognostic score which is widely used in bi…

Cited by 29SourcePDFScholar
2022

Unsupervised Learning of Equivariant Structure from Sequences

NeurIPS 2022accept

In this study, we present \textit{meta-sequential prediction} (MSP), an unsupervised framework to learn the symmetry from the time sequence of length at least three. Our method leverages the stationary property~(e.g. constant velocity, constant acceleration) of the time sequence to learn the underl…

2021

A General Class of Transfer Learning Regression without Implementation Cost

AAAI 2021technical

We propose a novel framework that unifies and extends existing methods of transfer learning (TL) for regression. To bridge a pretrained source model to the model on a target task, we introduce a density-ratio reweighting function, which is estimated through the Bayesian framework with a specific pri…

Cited by 9SourcePDFScholar
2020

Exchangeable Deep Neural Networks for Set-to-Set Matching and Learning

ECCV 2020poster

Matching two different sets of items, called heterogeneous set-to-set matching problem, has recently received attention as a promising problem. The difficulties are to extract features to match a correct pair of different sets and also preserve two types of exchangeability required for set-to-set ma…

Cited by 21SourcePDFScholar
2020

Robust Persistence Diagrams using Reproducing Kernels

NeurIPS 2020poster

Persistent homology has become an important tool for extracting geometric and topological features from data, whose multi-scale features are summarized in a persistence diagram. From a statistical perspective, however, persistence diagrams are very sensitive to perturbations in the input space. In t…

2019

Post Selection Inference with Incomplete Maximum Mean Discrepancy Estimator

ICLR 2019poster

Measuring divergence between two distributions is essential in machine learning and statistics and has various applications including binary classification, change point detection, and two-sample test. Furthermore, in the era of big data, designing divergence measure that is interpretable and can ha…

Cited by 28SourcePDFScholar
2019

Semi-flat minima and saddle points by embedding neural networks to overparameterization

NeurIPS 2019poster

We theoretically study the landscape of the training error for neural networks in overparameterized cases. We consider three basic methods for embedding a network into a wider one with more hidden units, and discuss whether a minimum point of the narrower network gives a minimum or saddle point of…

Cited by 31SourcePDFScholar
2018

Kernel Recursive ABC: Point Estimation with Intractable Likelihood

ICML 2018oral

We propose a novel approach to parameter estimation for simulator-based statistical models with intractable likelihood. Our proposed method involves recursive application of kernel ABC and kernel herding to the same observed data. We provide a theoretical explanation regarding why the approach works…

2018

Variational Learning on Aggregate Outputs with Gaussian Processes

NeurIPS 2018poster

While a typical supervised learning framework assumes that the inputs and the outputs are measured at the same levels of granularity, many applications, including global mapping of disease, only have access to outputs at a much coarser level than that of the inputs. Aggregation of outputs makes gene…

2017

A Linear-Time Kernel Goodness-of-Fit Test

NeurIPS 2017oral

We propose a novel adaptive test of goodness-of-fit, with computational cost linear in the number of samples. We learn the test features that best indicate the differences between observed samples and a reference model, by minimizing the false negative rate. These features are constructed via Stein'…

2016

Convergence guarantees for kernel-based quadrature rules in misspecified settings

NeurIPS 2016poster

Kernel-based quadrature rules are becoming important in machine learning and statistics, as they achieve super-$¥sqrt{n}$ convergence rates in numerical integration, and thus provide alternatives to Monte Carlo integration in challenging settings where integrands are expensive to evaluate or where i…

Cited by 57SourcePDFScholar
2016

Persistence weighted Gaussian kernel for topological data analysis

ICML 2016poster

Topological data analysis (TDA) is an emerging mathematical concept for characterizing shapes in complex data. In TDA, persistence diagrams are widely recognized as a useful descriptor of data, and can distinguish robust and noisy topological properties. This paper proposes a kernel method on persis…

Cited by 245SourcePDFScholar