← Search

Tomoharu Iwata

38 accepted papers

2025

Energy-consistent Neural Operators for Hamiltonian and Dissipative Partial Differential Equations

AISTATS 2025poster

The operator learning has received significant attention in recent years, with the aim of learning a mapping between function spaces. Prior works have proposed deep neural networks (DNNs) for learning such a mapping, enabling the learning of solution operators of partial differential equations (PDEs…

Cited by 0SourceScholar
2025

Hyperbolic PHATE: Visualizing Continuous Hierarchy of Latent Differentiation Structures

ICASSP 2025accepted

This paper proposes a method for embedding diffusion potentials into a hyperbolic space in order to visualize the differentiation structure consisting of diffusion and branching inherent in high-dimensional data. In recent years, the rapid development of single-cell sequencing in the field of biolog…

Cited by 0SourceScholar
2025

Importance-weighted Positive-unlabeled Learning for Distribution Shift Adaptation

AISTATS 2025oral

Positive and unlabeled (PU) learning is a fundamental task in many applications, which trains a binary classifier from only PU data. Existing PU learning methods typically assume that training and test distributions are identical. However, this assumption is often violated due to distribution shifts…

Cited by 0SourceScholar
2025

K$^2$IE: Kernel Method-based Kernel Intensity Estimators for Inhomogeneous Poisson Processes

ICML 2025poster

Kernel method-based intensity estimators, formulated within reproducing kernel Hilbert spaces (RKHSs), and classical kernel intensity estimators (KIEs) have been among the most easy-to-implement and feasible methods for estimating the intensity functions of inhomogeneous Poisson processes. While bot…

2025

Learning to Generate Projections for Reducing Dimensionality of Heterogeneous Linear Programming Problems

ICML 2025poster

We propose a data-driven method for reducing the dimensionality of linear programming problems (LPs) by generating instance-specific projection matrices using a neural network-based model. Once the model is trained using multiple LPs by maximizing the expected objective value, we can efficiently fin…

Cited by 0SourcePDFScholar
2025

Meta-learning Task-specific Regularization Weights for Few-shot Linear Regression

AISTATS 2025poster

We propose a few-shot learning method for linear regression, which learns how to choose regularization weights from multiple tasks with different feature spaces, and uses the knowledge for unseen tasks. Linear regression is ubiquitous in a wide variety of fields. Although regularization weight tunin…

Cited by 0SourceScholar
2025

Positive-Unlabeled Diffusion Models for Preventing Sensitive Data Generation

ICLR 2025poster

Diffusion models are powerful generative models but often generate sensitive data that are unwanted by users, mainly because the unlabeled training data frequently contain such sensitive data. Since labeling all sensitive data in the large-scale unlabeled training data is impractical, we address thi…

Cited by 0SourcePDFScholar
2025

Positive-unlabeled AUC Maximization under Covariate Shift

ICML 2025poster

Maximizing the area under the receiver operating characteristic curve (AUC) is a standard approach to imbalanced binary classification tasks. Existing AUC maximization methods typically assume that training and test distributions are identical. However, this assumption is often violated due to {\it…

Cited by 1SourcePDFScholar
2024

AUC Maximization under Positive Distribution Shift

NeurIPS 2024poster

Maximizing the area under the receiver operating characteristic curve (AUC) is a popular approach to imbalanced binary classification problems. Existing AUC maximization methods usually assume that training and test distributions are identical. However, this assumption is often violated in practice…

Cited by 0SourcePDFScholar
2024

Explanation-based Training with Differentiable Insertion/Deletion Metric-aware Regularizers

AISTATS 2024poster

The quality of explanations for the predictions made by complex machine learning predictors is often measured using insertion and deletion metrics, which assess the faithfulness of the explanations, i.e., how accurately the explanations reflect the predictor’s behavior. To improve the faithfulness,…

2024

Fast Iterative Hard Thresholding Methods with Pruning Gradient Computations

NeurIPS 2024poster

We accelerate the iterative hard thresholding (IHT) method, which finds \(k\) important elements from a parameter vector in a linear regression model. Although the plain IHT repeatedly updates the parameter vector during the optimization, computing gradients is the main bottleneck. Our method safely…

Cited by 0SourcePDFScholar
2024

Symplectic Neural Gaussian Processes for Meta-learning Hamiltonian Dynamics

IJCAI 2024poster

We propose a meta-learning method for modeling Hamiltonian dynamics from a limited number of data. Although Hamiltonian neural networks have been successfully used for modeling dynamics that obey the energy conservation law, they require many data to achieve high performance. The proposed method met…

2024

Warped Diffusion for Latent Differentiation Inference

AISTATS 2024poster

This paper proposes a Bayesian nonparametric diffusion model with a black-box warping function represented by a Gaussian process to infer potential diffusion structures latent in observed data, such as differentiation mechanisms of living cells and phylogenetic evolution processes of media informati…

2024

Zero-Shot Task Adaptation with Relevant Feature Information

AAAI 2024technical

We propose a method to learn prediction models such as classifiers for unseen target tasks where labeled and unlabeled data are absent but a few relevant input features for solving the tasks are given. Although machine learning requires data for training, data are often difficult to collect in pract…

2023

Meta-learning for Robust Anomaly Detection

AISTATS 2023poster

We propose a meta-learning method to improve the anomaly detection performance on unseen target tasks that have only unlabeled data. Existing meta-learning methods for anomaly detection have shown remarkable performance but require labeled data in target tasks. Although they can treat unlabeled data…

2022

Few-shot Learning for Feature Selection with Hilbert-Schmidt Independence Criterion

NeurIPS 2022accept

We propose a few-shot learning method for feature selection that can select relevant features given a small number of labeled instances. Existing methods require many labeled instances for accurate feature selection. However, sufficient instances are often unavailable. We use labeled instances in mu…

Cited by 11SourcePDFScholar
2022

Predictive variational Bayesian inference as risk-seeking optimization

AISTATS 2022poster

Since the Bayesian inference works poorly under model misspecification, various solutions have been explored to counteract the shortcomings. Recently proposed predictive Bayes (PB) that directly optimizes the Kullback Leibler divergence between the empirical distribution and the approximate predicti…

Cited by 3SourcePDFScholar
2022

Symplectic Spectrum Gaussian Processes: Learning Hamiltonians from Noisy and Sparse Data

NeurIPS 2022accept

Hamiltonian mechanics is a well-established theory for modeling the time evolution of systems with conserved quantities (called Hamiltonian), such as the total energy of the system. Recent works have parameterized the Hamiltonian by machine learning models (e.g., neural networks), allowing Hamiltoni…

Cited by 13SourcePDFScholar
2022

Tight Integration Of Neural- And Clustering-Based Diarization Through Deep Unfolding Of Infinite Gaussian Mixture Model

ICASSP 2022accepted

Speaker diarization has been investigated extensively as an important central task for meeting analysis. Recent trend shows that integration of end-to-end neural (EEND)- and clustering-based diarization is a promising approach to handle realistic conversational data containing overlapped speech with…

Cited by 0SourceScholar
2021

Loss function based second-order Jensen inequality and its application to particle variational inference

NeurIPS 2021poster

Bayesian model averaging, obtained as the expectation of a likelihood function by a posterior distribution, has been widely used for prediction, evaluation of uncertainty, and model selection. Various approaches have been developed to efficiently capture the information in the posterior distribution…

Cited by 6SourcePDFScholar
2020

Fast Deterministic CUR Matrix Decomposition with Accuracy Assurance

ICML 2020poster

The deterministic CUR matrix decomposition is a low-rank approximation method to analyze a data matrix. It has attracted considerable attention due to its high interpretability, which results from the fact that the decomposed matrices consist of subsets of the original columns and rows of the data m…

Cited by 14SourcePDFScholar
2020

Frame-Level Phoneme-Invariant Speaker Embedding for Text-Independent Speaker Recognition on Extremely Short Utterances

ICASSP 2020accepted

This paper investigates a phoneme-invariant speaker embedding approach for speaker recognition on extremely short utterances. Intuitively, phonemes are nuisance information for text-independent speaker recognition task since the contents of the speech are usually mismatched between enrolling and tes…

Cited by 0SourceScholar
2020

Reinforcement Learning in Latent Action Sequence Space

IROS 2020poster

One problem in real-world applications of reinforcement learning is the high dimensionality of the action search spaces, which comes from the combination of actions over time. To reduce the dimensionality of action sequence search spaces, macro actions have been studied, which are sequences of primi…

Cited by 5SourceScholar
2019

A Unified Framework for Feature-based Domain Adaptation of Neural Network Language Models

ICASSP 2019accepted

An important task for language models is the adaptation of general-domain models to specific target domains. For neural network-based language models, feature-based domain adaptation has been a popular method in previous research. Conventional methods use an adaptation feature providing context info…

Cited by 0SourceScholar
2019

Semi-supervised End-to-end Speech Recognition Using Text-to-speech and Autoencoders

ICASSP 2019accepted

We introduce speech and text autoencoders that share encoders and decoders with an automatic speech recognition (ASR) model to improve ASR performance with large speech only and text only training datasets. To build the speech and text autoencoders, we leverage state-of-the-art ASR and text-to-speec…

Cited by 44SourceScholar
2019

Spatially Aggregated Gaussian Processes with Multivariate Areal Outputs

NeurIPS 2019poster

We propose a probabilistic model for inferring the multivariate function from multiple areal data sets with various granularities. Here, the areal data are observed not at location points but at regions. Existing regression-based models can only utilize the sufficiently fine-grained auxiliary data s…

Cited by 33SourcePDFScholar
2019

Transfer Anomaly Detection by Inferring Latent Domain Representations

NeurIPS 2019poster

We propose a method to improve the anomaly detection performance on target domains by transferring knowledge on related domains. Although anomaly labels are valuable to learn anomaly detectors, they are difficult to obtain due to their rarity. To alleviate this problem, existing methods use anomalou…

Cited by 55SourcePDFScholar
2018

Language Model Domain Adaptation Via Recurrent Neural Networks with Domain-Shared and Domain-Specific Representations

ICASSP 2018accepted

Training recurrent neural network language models (RNNLMs) requires a large amount of data, which is difficult to collect for specific domains such as multiparty conversations. Data augmentation using external resources and model adaptation, which adjusts a model trained on a large amount of data to…

Cited by 0SourceScholar
2017

Localized Lasso for High-Dimensional Regression

AISTATS 2017poster

We introduce the localized Lasso, which learns models that both are interpretable and have a high predictive power in problems with high dimensionality d and small sample size n. More specifically, we consider a function defined by local sparse models, one at each data point. We introduce sample-wi…

Cited by 63SourcePDFScholar
2015

Cross-Domain Matching for Bag-of-Words Data via Kernel Embeddings of Latent Distributions

NeurIPS 2015poster

We propose a kernel-based method for finding matching between instances across different domains, such as multilingual documents and images with annotations. Each instance is assumed to be represented as a multiset of features, e.g., a bag-of-words representation for documents. The major difficulty…

Cited by 9SourcePDFScholar
2015

Cross-domain recommendation without shared users or items by sharing latent vector distributions

AISTATS 2015poster

We propose a cross-domain recommendation method for predicting the ratings of items in different domains, where neither users nor items are shared across domains. The proposed method is based on matrix factorization, which learns a latent vector for each user and each item. Matrix factorization tech…

Cited by 28SourcePDFScholar