← Search

Nhat Ho

88 accepted papers

2026

FACET: A Fragment-Aware Conformer Ensemble Transformer

ICLR 2026poster

Accurately predicting molecular properties requires effective integration of structural information from both 2D molecular graphs and their corresponding equilibrium conformer ensembles. In this work, we propose FACET, a scalable Structure-Aware Graph Transformer that efficiently aggregates features…

Cited by 0SourceScholar
2026

Fast Estimation of Wasserstein Distances via Regression on Sliced Wasserstein Distances

ICLR 2026poster

We address the problem of efficiently computing Wasserstein distances for multiple pairs of distributions drawn from a meta-distribution. To this end, we propose a fast estimation method based on regressing Wasserstein distance on sliced Wasserstein (SW) distances. Specifically, we leverage both sta…

Cited by 0SourcecodeScholar
2026

One-Prompt Strikes Back: Sparse Mixture of Experts for Prompt-based Continual Learning

ICLR 2026poster

Prompt-based methods have recently gained prominence in Continual Learning (CL) due to their strong performance and memory efficiency. A prevalent strategy in this paradigm assigns a dedicated subset of prompts to each task, which, while effective, incurs substantial computational overhead and cause…

Cited by 0SourcecodeScholar
2026

Revisit Visual Prompt Tuning: The Expressiveness of Prompt Experts

ICLR 2026poster

Visual Prompt Tuning (VPT) has proven effective for parameter-efficient adaptation of pre-trained vision models to downstream tasks by inserting task-specific learnable prompt tokens. Despite its empirical success, a comprehensive theoretical understanding of VPT remains an active area of research.…

Cited by 0SourcecodeScholar
2025

Beyond Losses Reweighting: Empowering Multi-Task Learning via the Generalization Perspective

ICCV 2025poster

Multi-task learning (MTL) trains deep neural networks to optimize several objectives simultaneously using a shared backbone, which leads to reduced computational costs, improved data efficiency, and enhanced performance through cross-task knowledge sharing. Although recent gradient manipulation tech…

Cited by 0SourcePDFScholar
2025

ExGra-Med: Extended Context Graph Alignment for Medical Vision-Language Models

NeurIPS 2025poster

State-of-the-art medical multi-modal LLMs (med-MLLMs), such as LLaVA-Med and BioMedGPT, primarily depend on scaling model size and data volume, with training driven largely by autoregressive objectives. However, we reveal that this approach can lead to weak vision-language alignment, making these mo…

Cited by 0SourceScholar
2025

Improving Generalization with Flat Hilbert Bayesian Inference

ICML 2025poster

We introduce Flat Hilbert Bayesian Inference (FHBI), an algorithm designed to enhance generalization in Bayesian inference. Our approach involves an iterative two-step procedure with an adversarial functional perturbation step and a functional descent step within the reproducing kernel Hilbert space…

Cited by 0SourcePDFScholar
2025

Lightspeed Geometric Dataset Distance via Sliced Optimal Transport

ICML 2025poster

We introduce sliced optimal transport dataset distance (s-OTDD), a model-agnostic, embedding-agnostic approach for dataset comparison that requires no training, is robust to variations in the number of classes, and can handle disjoint label sets. The core innovation is Moment Transform Projection…

2025

On Minimax Estimation of Parameters in Softmax-Contaminated Mixture of Experts

NeurIPS 2025poster

The softmax-contaminated mixture of experts (MoE) model is deployed when a large-scale pre-trained model, which plays the role of a fixed expert, is fine-tuned for learning downstream tasks by including a new contamination part, or prompt, functioning as a new, trainable expert. Despite its populari…

Cited by 0SourceScholar
2025

On Zero-Initialized Attention: Optimal Prompt and Gating Factor Estimation

ICML 2025poster

LLaMA-Adapter has recently emerged as an efficient fine-tuning technique for LLaMA models, leveraging zero-initialized attention to stabilize training and enhance performance. However, despite its empirical success, the theoretical foundations of zero-initialized attention remain largely unexplored.…

Cited by 1SourcePDFScholar
2025

RepLoRA: Reparameterizing Low-rank Adaptation via the Perspective of Mixture of Experts

ICML 2025poster

Low-rank Adaptation (LoRA) has emerged as a powerful and efficient method for fine-tuning large-scale foundation models. Despite its popularity, the theoretical understanding of LoRA has remained underexplored. In this paper, we present a theoretical analysis of LoRA by examining its connection to t…

Cited by 1SourcePDFScholar
2025

Revisiting Prefix-tuning: Statistical Benefits of Reparameterization among Prompts

ICLR 2025poster

Prompt-based techniques, such as prompt-tuning and prefix-tuning, have gained prominence for their efficiency in fine-tuning large pre-trained models. Despite their widespread adoption, the theoretical foundations of these methods remain limited. For instance, in prefix-tuning, we observe that a key…

Cited by 4SourcePDFScholar
2025

Statistical Advantages of Perturbing Cosine Router in Mixture of Experts

ICLR 2025poster

The cosine router in Mixture of Experts (MoE) has recently emerged as an attractive alternative to the conventional linear router. Indeed, the cosine router demonstrates favorable performance in image and language tasks and exhibits better ability to mitigate the representation collapse issue, which…

Cited by 6SourcePDFScholar
2025

Understanding Expert Structures on Minimax Parameter Estimation in Contaminated Mixture of Experts

AISTATS 2025poster

We conduct the convergence analysis of parameter estimation in the contaminated mixture of experts. This model is motivated from the prompt learning problem where ones utilize prompts, which can be formulated as experts, to fine-tune a large-scale pre-trained model for learning downstream tasks. The…

Cited by 0SourceScholar
2025

X-Drive: Cross-modality Consistent Multi-Sensor Data Synthesis for Driving Scenarios

ICLR 2025poster

Recent advancements have exploited diffusion models for the synthesis of either LiDAR point clouds or camera image data in driving scenarios. Despite their success in modeling single-modality data marginal distribution, there is an under- exploration in the mutual reliance between different modaliti…

2024

A Bayesian Approach for Personalized Federated Learning in Heterogeneous Settings

NeurIPS 2024poster

Federated learning (FL), through its privacy-preserving collaborative learning approach, has significantly empowered decentralized devices. However, constraints in either data and/or computational resources among participating clients introduce several challenges in learning, including the inabilit…

Cited by 0SourcePDFScholar
2024

A General Theory for Softmax Gating Multinomial Logistic Mixture of Experts

ICML 2024poster

Mixture-of-experts (MoE) model incorporates the power of multiple submodels via gating functions to achieve greater performance in numerous regression and classification applications. From a theoretical perspective, while there have been previous attempts to comprehend the behavior of that model und…

Cited by 8SourcePDFScholar
2024

Bayesian Nonparametrics Meets Data-Driven Distributionally Robust Optimization

NeurIPS 2024poster

Training machine learning and statistical models often involves optimizing a data-driven risk criterion. The risk is usually computed with respect to the empirical data distribution, but this may result in poor and unstable out-of-sample performance due to distributional uncertainty. In the spirit o…

2024

Beyond Vanilla Variational Autoencoders: Detecting Posterior Collapse in Conditional and Hierarchical Variational Autoencoders

ICLR 2024poster

The posterior collapse phenomenon in variational autoencoder (VAE), where the variational posterior distribution closely matches the prior distribution, can hinder the quality of the learned latent variables. As a consequence of posterior collapse, the latent variables extracted by the encoder in VA…

Cited by 3SourcePDFScholar
2024

Diffeomorphic Mesh Deformation via Efficient Optimal Transport for Cortical Surface Reconstruction

ICLR 2024poster

Mesh deformation plays a pivotal role in many 3D vision tasks including dynamic simulations, rendering, and reconstruction. However, defining an efficient discrepancy between predicted and target meshes remains an open problem. A prevalent approach in current deep learning is the set-based approach…

Cited by 1SourcePDFScholar
2024

Fast Approximation of the Generalized Sliced-Wasserstein Distance

ICASSP 2024accepted

Generalized sliced-Wasserstein distance is a variant of slicedWasserstein distance that exploits the power of non-linear projection through a given defining function to better capture the complex structures of probability distributions. Similar to the sliced-Wasserstein distance, generalized slicedW…

Cited by 0SourceScholar
2024

FuseMoE: Mixture-of-Experts Transformers for Fleximodal Fusion

NeurIPS 2024poster

As machine learning models in critical fields increasingly grapple with multimodal data, they face the dual challenges of handling a wide array of modalities, often incomplete due to missing elements, and the temporal irregularity and sparsity of collected samples. Successfully leveraging this compl…

Cited by 22SourcePDFScholar
2024

Hierarchical Hybrid Sliced Wasserstein: A Scalable Metric for Heterogeneous Joint Distributions

NeurIPS 2024poster

Sliced Wasserstein (SW) and Generalized Sliced Wasserstein (GSW) have been widely used in applications due to their computational and statistical scalability. However, the SW and the GSW are only defined between distributions supported on a homogeneous domain. This limitation prevents their usage in…

2024

Improving Computational Complexity in Statistical Models with Local Curvature Information

ICML 2024poster

It is known that when the statistical models are singular, i.e., the Fisher information matrix at the true parameter is degenerate, the fixed step-size gradient descent algorithm takes polynomial number of steps in terms of the sample size $n$ to converge to a final statistical radius around the tru…

Cited by 0SourcePDFScholar
2024

Integrating Efficient Optimal Transport and Functional Maps For Unsupervised Shape Correspondence Learning

CVPR 2024poster

In the realm of computer vision and graphics accurately establishing correspondences between geometric 3D shapes is pivotal for applications like object tracking registration texture transfer and statistical shape analysis. Moving beyond traditional hand-crafted and data-driven feature learning meth…

Cited by 4SourcePDFScholar
2024

Mixture of Experts Meets Prompt-Based Continual Learning

NeurIPS 2024poster

Exploiting the power of pre-trained models, prompt-based approaches stand out compared to other continual learning solutions in effectively preventing catastrophic forgetting, even with very few learnable parameters and without the need for a memory buffer. While existing prompt-based continual lear…

2024

Neural Collapse for Cross-entropy Class-Imbalanced Learning with Unconstrained ReLU Features Model

ICML 2024poster

The current paradigm of training deep neural networks for classification tasks includes minimizing the empirical risk, pushing the training loss value towards zero even after the training classification error has vanished. In this terminal phase of training, it has been observed that the last-layer…

Cited by 12SourcePDFScholar
2024

Revisiting Deep Audio-Text Retrieval Through the Lens of Transportation

ICLR 2024poster

The Learning-to-match (LTM) framework proves to be an effective inverse optimal transport approach for learning the underlying ground metric between two sources of data, facilitating subsequent matching. However, the conventional LTM framework faces scalability challenges, necessitating the use of t…

2024

Sigmoid Gating is More Sample Efficient than Softmax Gating in Mixture of Experts

NeurIPS 2024poster

The softmax gating function is arguably the most popular choice in mixture of experts modeling. Despite its widespread use in practice, the softmax gating may lead to unnecessary competition among experts, potentially causing the undesirable phenomenon of representation collapse due to its inherent…

Cited by 7SourcePDFScholar
2024

Statistical Perspective of Top-K Sparse Softmax Gating Mixture of Experts

ICLR 2024poster

Top-K sparse softmax gating mixture of experts has been widely used for scaling up massive deep-learning architectures without increasing the computational cost. Despite its popularity in real-world applications, the theoretical understanding of that gating function has remained an open problem. The…

Cited by 17SourcePDFScholar
2024

Structure-Aware E(3)-Invariant Molecular Conformer Aggregation Networks

ICML 2024poster

A molecule’s 2D representation consists of its atoms, their attributes, and the molecule’s covalent bonds. A 3D (geometric) representation of a molecule is called a conformer and consists of its atom types and Cartesian coordinates. Every conformer has a potential energy, and the lower this energy,…

2024

Towards Convergence Rates for Parameter Estimation in Gaussian-gated Mixture of Experts

AISTATS 2024poster

Originally introduced as a neural network for ensemble learning, mixture of experts (MoE) has recently become a fundamental building block of highly successful modern deep neural networks for heterogeneous data analysis in several applications of machine learning and statistics. Despite its populari…

2023

A Primal-Dual Framework for Transformers and Neural Networks

ICLR 2023top-25%

Self-attention is key to the remarkable success of transformers in sequence modeling tasks including many applications in natural language processing and computer vision. Like neural network layers, these attention mechanisms are often developed by heuristics and experience. To provide a principled…

Cited by 18SourcePDFScholar
2023

A Probabilistic Framework for Pruning Transformers Via a Finite Admixture of Keys

ICASSP 2023accepted

Pairwise dot product-based self-attention is key to the success of transformers which achieve state-of-the-art performance across a variety of applications in language and vision, but are costly to compute. It has been shown that most attention scores and keys in transformers are redundant and can b…

Cited by 0SourceScholar
2023

Designing Robust Transformers using Robust Kernel Density Estimation

NeurIPS 2023poster

Transformer-based architectures have recently exhibited remarkable successes across different domains beyond just powering large language models. However, existing approaches typically focus on predictive accuracy and computational cost, largely ignoring certain other practical issues such as robust…

Cited by 9SourcePDFScholar
2023

Global-Local Regularization Via Distributional Robustness

AISTATS 2023poster

Despite superior performance in many situations, deep neural networks are often vulnerable to adversarial examples and distribution shifts, limiting model generalization ability in real-world applications. To alleviate these problems, recent approaches leverage distributional robustness optimization…

2023

Hierarchical Sliced Wasserstein Distance

ICLR 2023poster

Sliced Wasserstein (SW) distance has been widely used in different application scenarios since it can be scaled to a large number of supports without suffering from the curse of dimensionality. The value of sliced Wasserstein distance is the average of transportation cost between one-dimensional rep…

2023

Joint Self-Supervised Image-Volume Representation Learning with Intra-inter Contrastive Clustering

AAAI 2023technical

Collecting large-scale medical datasets with fully annotated samples for training of deep networks is prohibitively expensive, especially for 3D volume data. Recent breakthroughs in self-supervised learning (SSL) offer the ability to overcome the lack of labeled training samples by learning feature…

Cited by 23SourcePDFScholar
2023

LVM-Med: Learning Large-Scale Self-Supervised Vision Models for Medical Imaging via Second-order Graph Matching

NeurIPS 2023poster

Obtaining large pre-trained models that can be fine-tuned to new tasks with limited annotated samples has remained an open challenge for medical imaging data. While pre-trained networks on ImageNet and vision-language foundation models trained on web-scale data are the prevailing approaches, their e…

2023

Markovian Sliced Wasserstein Distances: Beyond Independent Projections

NeurIPS 2023poster

Sliced Wasserstein (SW) distance suffers from redundant projections due to independent uniform random projecting directions. To partially overcome the issue, max K sliced Wasserstein (Max-K-SW) distance ($K\geq 1$), seeks the best discriminative orthogonal projecting directions. Despite being able…

2023

Minimax Optimal Rate for Parameter Estimation in Multivariate Deviated Models

NeurIPS 2023poster

We study the maximum likelihood estimation (MLE) in the multivariate deviated model where the data are generated from the density function $(1-\lambda^{\ast})h_{0}(x)+\lambda^{\ast}f(x|\mu^{\ast}, \Sigma^{\ast})$ in which $h_{0}$ is a known function, $\lambda^{\ast} \in [0,1]$ and $(\mu^{\ast}, \Sig…

Cited by 4SourcePDFScholar
2023

Neural Collapse in Deep Linear Networks: From Balanced to Imbalanced Data

ICML 2023poster

Modern deep neural networks have achieved impressive performance on tasks from image classification to natural language processing. Surprisingly, these complex systems with massive amounts of parameters exhibit the same structural properties in their last-layer features and classifiers across canoni…

2023

On Cross-Layer Alignment for Model Fusion of Heterogeneous Neural Networks

ICASSP 2023accepted

OTFusion, or layer-wise model fusion via optimal transport, applies soft neuron association to unify different pre-trained networks. Despite its effectiveness in saving computational resources, OTFusion requires the input networks to have the same number of layers. To address this issue, we propose…

Cited by 0SourceScholar
2023

On Excess Mass Behavior in Gaussian Mixture Models with Orlicz-Wasserstein Distances

ICML 2023poster

Dirichlet Process mixture models (DPMM) in combination with Gaussian kernels have been an important modeling tool for numerous data domains arising from biological, physical, and social sciences. However, this versatility in applications does not extend to strong theoretical guarantees for the under…

Cited by 5SourcePDFScholar
2023

Revisiting Over-smoothing and Over-squashing Using Ollivier-Ricci Curvature

ICML 2023poster

Graph Neural Networks (GNNs) had been demonstrated to be inherently susceptible to the problems of over-smoothing and over-squashing. These issues prohibit the ability of GNNs to model complex graph interactions by limiting their effectiveness in taking into account distant information. Our study re…

2023

Self-Attention Amortized Distributional Projection Optimization for Sliced Wasserstein Point-Cloud Reconstruction

ICML 2023poster

Max sliced Wasserstein (Max-SW) distance has been widely known as a solution for less discriminative projections of sliced Wasserstein (SW) distance. In applications that have various independent pairs of probability measures, amortized projection optimization is utilized to predict the ``max" proje…

2022

Entropic Gromov-Wasserstein between Gaussian Distributions

ICML 2022spotlight

We study the entropic Gromov-Wasserstein and its unbalanced version between (unbalanced) Gaussian distributions with different dimensions. When the metric is the inner product, which we refer to as inner product Gromov-Wasserstein (IGW), we demonstrate that the optimal transportation plans of entrop…

2022

FourierFormer: Transformer Meets Generalized Fourier Integral Theorem

NeurIPS 2022accept

Multi-head attention empowers the recent success of transformers, the state-of-the-art models that have achieved remarkable success in sequence modeling and beyond. These attention mechanisms compute the pairwise dot products between the queries and keys, which results from the use of unnormalized G…

Cited by 41SourcePDFScholar
2022

Improving Mini-batch Optimal Transport via Partial Transportation

ICML 2022spotlight

Mini-batch optimal transport (m-OT) has been widely used recently to deal with the memory issue of OT in large-scale applications. Despite their practicality, m-OT suffers from misspecified mappings, namely, mappings that are optimal on the mini-batch level but are partially wrong in the comparison…

Cited by 50SourcePDFScholar
2022

Improving Transformer with an Admixture of Attention Heads

NeurIPS 2022accept

Transformers with multi-head self-attention have achieved remarkable success in sequence modeling and beyond. However, they suffer from high computational and memory complexities for computing the attention matrix at each head. Recently, it has been shown that those attention matrices lie on a low-d…

Cited by 29SourcePDFScholar
2022

Improving Transformers with Probabilistic Attention Keys

ICML 2022spotlight

Multi-head attention is a driving force behind state-of-the-art transformers, which achieve remarkable performance across a variety of natural language processing (NLP) and computer vision tasks. It has been observed that for many applications, those attention heads learn redundant embedding, and mo…

2022

On Multimarginal Partial Optimal Transport: Equivalent Forms and Computational Complexity

AISTATS 2022poster

We study the multi-marginal partial optimal transport (POT) problem between $m$ discrete (unbalanced) measures with at most $n$ supports. We first prove that we can obtain two equivalent forms of the multimarginal POT problem in terms of the multimarginal optimal transport problem via novel extensio…

Cited by 13SourcePDFScholar
2022

On Structured Filtering-Clustering: Global Error Bound and Optimal First-Order Algorithms

AISTATS 2022poster

The filtering-clustering models, including trend filtering and convex clustering, have become an important source of ideas and modeling tools in machine learning and related fields. The statistical guarantee of optimal solutions in these models has been extensively studied yet the investigations on…

Cited by 2SourcePDFScholar
2022

On Transportation of Mini-batches: A Hierarchical Approach

ICML 2022spotlight

Mini-batch optimal transport (m-OT) has been successfully used in practical applications that involve probability measures with a very high number of supports. The m-OT solves several smaller optimal transport problems and then returns the average of their costs and transportation plans. Despite its…

Cited by 21SourcePDFScholar
2022

Refined Convergence Rates for Maximum Likelihood Estimation under Finite Mixture Models

ICML 2022oral

We revisit the classical problem of deriving convergence rates for the maximum likelihood estimator (MLE) in finite mixture models. The Wasserstein distance has become a standard loss function for the analysis of parameter estimation in these models, due in part to its ability to circumvent label sw…

2022

Stochastic Multiple Target Sampling Gradient Descent

NeurIPS 2022accept

Sampling from an unnormalized target distribution is an essential problem with many applications in probabilistic inference. Stein Variational Gradient Descent (SVGD) has been shown to be a powerful method that iteratively updates a set of particles to approximate the distribution of interest. Furth…

2022

Towards Statistical and Computational Complexities of Polyak Step Size Gradient Descent

AISTATS 2022poster

We study the statistical and computational complexities of the Polyak step size gradient descent algorithm under generalized smoothness and {Ł}ojasiewicz conditions of the population loss function, namely, the limit of the empirical loss function when the sample size goes to infinity, and the stabil…

Cited by 10SourcePDFScholar
2022

Weak Separation in Mixture Models and Implications for Principal Stratification

AISTATS 2022poster

Principal stratification is a popular framework for addressing post-randomization complications, often in conjunction with finite mixture models for estimating the causal effects of interest. Unfortunately, standard estimators of mixture parameters, like the MLE, are known to exhibit pathological be…

Cited by 16SourcePDFScholar
2021

Distributional Sliced-Wasserstein and Applications to Generative Modeling

ICLR 2021spotlight

Sliced-Wasserstein distance (SW) and its variant, Max Sliced-Wasserstein distance (Max-SW), have been used widely in the recent years due to their fast computation and scalability even when the probability measures lie in a very high dimensional space. However, SW requires many unnecessary projectio…

2021

Flow-based Alignment Approaches for Probability Measures in Different Spaces

AISTATS 2021poster

Gromov-Wasserstein (GW) is a powerful tool to compare probability measures whose supports are in different metric spaces. However, GW suffers from a computational drawback since it requires to solve a complex non-convex quadratic program. In this work, we consider a specific family of cost metrics,…

2021

Improving Relational Regularized Autoencoders with Spherical Sliced Fused Gromov Wasserstein

ICLR 2021poster

Relational regularized autoencoder (RAE) is a framework to learn the distribution of data by minimizing a reconstruction loss together with a relational regularization on the prior of latent space. A recent attempt to reduce the inner discrepancy between the prior and aggregated posterior distributi…

Cited by 31SourcePDFScholar
2021

On Robust Optimal Transport: Computational Complexity and Barycenter Computation

NeurIPS 2021poster

We consider robust variants of the standard optimal transport, named robust optimal transport, where marginal constraints are relaxed via Kullback-Leibler divergence. We show that Sinkhorn-based algorithms can approximate the optimal cost of robust optimal transport in $\widetilde{\mathcal{O}}(\frac…

Cited by 48SourcePDFScholar
2021

On the Minimax Optimality of the EM Algorithm for Learning Two-Component Mixed Linear Regression

AISTATS 2021poster

We study the convergence rates of the EM algorithm for learning two-component mixed linear regression under all regimes of signal-to-noise ratio (SNR). We resolve a long-standing question that many recent results have attempted to tackle: we completely characterize the convergence behavior of EM, an…

Cited by 49SourcePDFScholar
2021

Point-Set Distances for Learning Representations of 3D Point Clouds

ICCV 2021poster

Learning an effective representation of 3D point clouds requires a good metric to measure the discrepancy between two 3D point sets, which is non-trivial due to their irregularity. Most of the previous works resort to using the Chamfer discrepancy or Earth Mover's distance, but those metrics are eit…

Cited by 94PDFcodeScholar
2021

Structured Dropout Variational Inference for Bayesian Neural Networks

NeurIPS 2021poster

Approximate inference in Bayesian deep networks exhibits a dilemma of how to yield high fidelity posterior approximations while maintaining computational efficiency and scalability. We tackle this challenge by introducing a novel variational structured approximation inspired by the Bayesian interpre…

Cited by 10SourcePDFScholar
2020

Fast Algorithms for Computational Optimal Transport and Wasserstein Barycenter

AISTATS 2020poster

We provide theoretical complexity analysis for new algorithms to compute the optimal transport (OT) distance between two discrete probability distributions, and demonstrate their favorable practical performance compared to state-of-art primal-dual algorithms. First, we introduce the \emph{accelerate…

Cited by 42SourcePDFScholar
2020

Fixed-Support Wasserstein Barycenters: Computational Hardness and Fast Algorithm

NeurIPS 2020poster

We study the fixed-support Wasserstein barycenter problem (FS-WBP), which consists in computing the Wasserstein barycenter of $m$ discrete probability measures supported on a finite metric space of size $n$. We show first that the constraint matrix arising from the standard linear programming (LP) r…

Cited by 65SourcePDFScholar
2020

On Unbalanced Optimal Transport: An Analysis of Sinkhorn Algorithm

ICML 2020poster

We provide a computational complexity analysis for the Sinkhorn algorithm that solves the entropic regularized Unbalanced Optimal Transport (UOT) problem between two measures of possibly different masses with at most $n$ components. We show that the complexity of the Sinkhorn algorithm for finding a…

Cited by 112SourcePDFScholar
2020

Projection Robust Wasserstein Distance and Riemannian Optimization

NeurIPS 2020spotlight

Projection robust Wasserstein (PRW) distance, or Wasserstein projection pursuit (WPP), is a robust variant of the Wasserstein distance. Recent work suggests that this quantity is more robust than the standard Wasserstein distance, in particular when comparing probability measures in high-dimensions.…

2020

Sharp Analysis of Expectation-Maximization for Weakly Identifiable Models

AISTATS 2020poster

We study a class of weakly identifiable location-scale mixture models for which the maximum likelihood estimates based on $n$ i.i.d. samples are known to have lower accuracy than the classical $n^{- \frac{1}{2}}$ error. We investigate whether the Expectation-Maximization (EM) algorithm also converge…

Cited by 33SourcePDFScholar
2019

Probabilistic Multilevel Clustering via Composite Transportation Distance

AISTATS 2019poster

We propose a novel probabilistic approach to multilevel clustering problems based on composite transportation distance, which is a variant of transportation distance where the underlying metric is Kullback-Leibler divergence. Our method involves solving a joint optimization problem over spaces of pr…

Cited by 26SourcePDFScholar
2017

Multilevel Clustering via Wasserstein Means

ICML 2017poster

We propose a novel approach to the problem of multilevel clustering, which aims to simultaneously partition data in each group and discover grouping patterns among groups in a potentially large hierarchically structured corpus of data. Our method involves a joint optimization formulation over severa…