← Search

Masashi Sugiyama

179 accepted papers

2026

Accessible, Realistic, and Fair Evaluation of Positive-Unlabeled Learning Algorithms

ICLR 2026poster

Positive-unlabeled (PU) learning is a weakly supervised binary classification problem, in which the goal is to learn a binary classifier from only positive and unlabeled data, without access to negative data. In recent years, many PU learning algorithms have been developed to improve model performan…

Cited by 0SourceScholar
2026

Bilateral Information-aware Test-time Adaptation for Vision-Language Models

ICLR 2026poster

Test-time adaptation (TTA) fine-tunes models using new data encountered during inference, which enables the vision-language models to handle test data with covariant shifts. Unlike training-time adaptation, TTA does not require a test-distributed validation set or consider the worst-case distributio…

Cited by 0SourcecodeScholar
2026

CARPRT: Class-Aware Zero-Shot Prompt Reweighting for Vision-Language Model

ICLR 2026poster

Pre-trained vision-language models (VLMs) enable zero-shot image classification by computing the similarity score between an image and textual descriptions, typically formed by inserting a class label (e.g., "cat") into a prompt (e.g., "a photo of a").Existing studies have shown that the score betwe…

Cited by 0SourceScholar
2026

Decomposing the Basic Abilities of Large Language Models: Mitigating Cross-Task Interference in Multi-Task Instruct-Tuning

ICML 2026poster

Recently, the prominent performance of large language models (LLMs) has been largely driven by multi-task instruct-tuning. Unfortunately, this training paradigm suffers from a key issue, named cross-task interference, due to conflicting gradients over shared parameters among different tasks. Some pr…

Cited by 0SourceScholar
2026

Decoupling the Class Label and the Target Concept in Machine Unlearning

ICLR 2026poster

Machine unlearning as an emerging research topic for data regulations, aims to adjust a trained model to approximate a retrained one that excludes a portion of training data. Previous studies showed that class-wise unlearning is effective in forgetting the knowledge of a training class, either throu…

Cited by 0SourceScholar
2026

Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards

ICML 2026poster

Reinforcement Learning from Human Feedback (RLFH) or Verifiable Rewards (RLVR) are two key steps in the post-training of modern Language Models (LMs). A common problem is reward hacking, where the policy may exploit inaccuracies of the reward and learn an unintended behavior. Most previous works add…

Cited by 0SourceScholar
2026

Incorporating Importance Weighting in Optimal Transport Based Domain Alignment

ICML 2026poster

Domain adaptation theory studies upper bounds on the target risk in order to mitigate performance loss of machine learning models due to distribution shift. In this paper, we take a closer look at the optimization of one such bound based on optimal transport (OT) and propose various strategies that …

Cited by 0SourceScholar
2026

Learning from Label Proportions via Proportional Value Classification

ICLR 2026poster

Learning from Label Proportions~(LLP) aims to use bags of instances associated with the proportions of each label within the bag to learn an instance-level classifier. Proportion matching is a widely used strategy that aligns the average model outputs of all instances in a bag with the label proport…

Cited by 0SourcecodeScholar
2026

Positive-Unlabeled Learning with Extreme Scarcity of Labeled Positives

ICML 2026poster

Positive-Unlabeled (PU) learning is a weakly-supervised paradigm that trains a binary classifier from labeled positive and unlabeled instances. In PU risk estimation, the empirical risk consists of an unlabeled term and a positive term. In this paper, we observe that when labeled positives are scarc…

Cited by 0SourceScholar
2026

Positive–Unlabeled Reinforcement Learning Distillation for On-Premise Small Models

ICML 2026poster

Due to constraints on privacy, cost, and latency, on-premise deployment of small models is increasingly common. However, most practical pipelines stop at supervised fine-tuning (SFT) and fail to reach the reinforcement learning (RL) alignment stage. The main reason is that RL alignment typically req…

Cited by 0SourceScholar
2026

Practical estimation of the optimal classification error with soft labels and calibration

ICLR 2026poster

While the performance of machine learning systems has experienced significant improvement in recent years, relatively little attention has been paid to the fundamental question: to what extent can we improve our models? This paper provides a means of answering this question in the setting of binary…

Cited by 0SourcecodeScholar
2026

Proteo-R1: Thinking Foundation Models for De Novo Protein Binder Design

ICML 2026poster

Recent advances in generative diffusion and flow-matching models have revolutionized molecular design, enabling the creation of novel proteins, small molecules, and RNA sequences with unprecedented fidelity. Yet, these models remain intuitive rather than intelligent—they generate without reasoning. …

Cited by 0SourceScholar
2026

Reinforcement Learning from Bagged Reward

ICML 2026poster

In Reinforcement Learning (RL), it is commonly assumed that an immediate reward signal is generated for each action taken by the agent, helping the agent maximize cumulative rewards to obtain the optimal policy. However, in many real-world scenarios, designing immediate reward signals is difficult; …

Cited by 0SourceScholar
2026

Rethinking Consistent Multi-Label Classification under Inexact Supervision

ICLR 2026poster

Partial multi-label learning and complementary multi-label learning are two popular weakly supervised multi-label classification paradigms that aim to alleviate the high annotation costs of collecting precisely annotated multi-label data. In partial multi-label learning, each instance is annotated w…

Cited by 0SourceScholar
2026

Robust Learning from Noisily Labeled Long-Tailed Data via Fairness Regularizer

AAAI 2026technical

Both long-tailed and noisily labeled data frequently appear in real-world applications and impose significant challenges for learning. Most prior works treat either problem in an isolated way and do not explicitly consider the coupling effects of the two. Our empirical observation reveals that such

Cited by 0SourcePDFScholar
2026

Towards Understanding Valuable Preference Data for Large Language Model Alignment

ICLR 2026poster

Large language model (LLM) alignment is typically achieved through learning from human preference comparisons, making the quality of preference data critical to its success. Existing studies often pre-process raw training datasets to identify valuable preference pairs using external reward models or…

Cited by 0SourceScholar
2026

Unlocking the Power of Co-Occurrence in CLIP: A DualPrompt-Driven Method for Training-Free Zero-Shot Multi-Label Classification

ICLR 2026poster

Contrastive Language-Image Pretraining (CLIP) has exhibited powerful zero-shot capacity in various single-label image classification tasks. However, when applying to the multi-label scenarios, CLIP suffers from significant performance declines due to the lack of explicit exploitation of co-occurrenc…

Cited by 0SourceScholar
2025

Action-Agnostic Point-Level Supervision for Temporal Action Detection

AAAI 2025technical

We propose action-agnostic point-level (AAPL) supervision for temporal action detection to achieve accurate action instance detection with a lightly annotated dataset. In the proposed scheme, a small portion of video frames is sampled in an unsupervised manner and presented to human annotators, who…

2025

Adaptive Localization of Knowledge Negation for Continual LLM Unlearning

ICML 2025poster

With the growing deployment of large language models (LLMs) across diverse domains, concerns regarding their safety have grown substantially. LLM unlearning has emerged as a pivotal approach to removing harmful or unlawful contents while maintaining utility. Despite increasing interest, the challeng…

Cited by 0SourcePDFScholar
2025

Domain Adaptation and Entanglement: an Optimal Transport Perspective

AISTATS 2025poster

Current machine learning systems are brittle in the face of distribution shifts (DS), where the target distribution that the system is tested on differs from the source distribution used to train the system. This problem of robustness to DS has been studied extensively in the field of domain adaptat…

Cited by 0SourceScholar
2025

Generalized Linear Bandits: Almost Optimal Regret with One-Pass Update

NeurIPS 2025poster

We study the generalized linear bandit (GLB) problem, a contextual multi-armed bandit framework that extends the classical linear model by incorporating a non-linear link function, thereby modeling a broad class of reward distributions such as Bernoulli and Poisson. While GLBs are widely applicable…

Cited by 0SourceScholar
2025

Label Distribution Learning with Biased Annotations Assisted by Multi-Label Learning

IJCAI 2025

Multi-label learning (MLL) has gained attention for its ability to represent real-world data. Label Distribution Learning (LDL), an extension of MLL to learning from label distributions, faces challenges in collecting accurate label distributions. To address the issue of biased annotations, based on

Cited by 0SourcePDFScholar
2025

Learning View-invariant World Models for Visual Robotic Manipulation

ICLR 2025poster

Robotic manipulation tasks often rely on visual inputs from cameras to perceive the environment. However, previous approaches still suffer from performance degradation when the camera’s viewpoint changes during manipulation. In this paper, we propose ReViWo (Representation learning for View-invarian…

Cited by 0SourcePDFScholar
2025

Non-stationary Online Learning for Curved Losses: Improved Dynamic Regret via Mixability

ICML 2025poster

Non-stationary online learning has drawn much attention in recent years. Despite considerable progress, dynamic regret minimization has primarily focused on convex functions, leaving the functions with stronger curvature (e.g., squared or logistic loss) underexplored. In this work, we address this g…

Cited by 0SourcePDFScholar
2025

Realistic Evaluation of Deep Partial-Label Learning Algorithms

ICLR 2025spotlight

Partial-label learning (PLL) is a weakly supervised learning problem in which each example is associated with multiple candidate labels and only one is the true label. In recent years, many deep PLL algorithms have been developed to improve model performance. However, we find that some early develop…

Cited by 1SourcePDFScholar
2025

Robust Multi-View Learning via Representation Fusion of Sample-Level Attention and Alignment of Simulated Perturbation

ICCV 2025poster

Recently, multi-view learning (MVL) has garnered significant attention due to its ability to fuse discriminative information from multiple views. However, real-world multi-view datasets are often heterogeneous and imperfect, which usually causes MVL methods designed for specific combinations of view…

2025

Sharpness-Aware Black-Box Optimization

ICLR 2025poster

Black-box optimization algorithms have been widely used in various machine learning problems, including reinforcement learning and prompt fine-tuning. However, directly optimizing the training loss value, as commonly done in existing black-box optimization methods, could lead to suboptimal model qua…

Cited by 0SourcePDFScholar
2025

Towards Effective Evaluations and Comparisons for LLM Unlearning Methods

ICLR 2025poster

The imperative to eliminate undesirable data memorization underscores the significance of machine unlearning for large language models (LLMs). Recent research has introduced a series of promising unlearning methods, notably boosting the practical significance of the field. Nevertheless, adopting a p…

Cited by 0SourcePDFScholar
2025

Towards Out-of-Modal Generalization without Instance-level Modal Correspondence

ICLR 2025poster

The world is understood from various modalities, such as appearance, sound, language, etc. Since each modality only partially represents objects in a certain physical meaning, leveraging additional ones is beneficial in both theory and practice. However, exploiting novel modalities normally requires…

Cited by 1SourcePDFScholar
2024

A General Framework for Learning from Weak Supervision

ICML 2024poster

Weakly supervised learning generally faces challenges in applicability to various scenarios with diverse weak supervision and in scalability due to the complexity of existing algorithms, thereby hindering the practical deployment. This paper introduces a general framework for learning from weak supe…

2024

Accurate Forgetting for Heterogeneous Federated Continual Learning

ICLR 2024poster

Recent years have witnessed a burgeoning interest in federated learning (FL). However, the contexts in which clients engage in sequential learning remain under- explored. Bridging FL and continual learning (CL) gives rise to a challenging practical problem: federated continual learning (FCL). Existi…

2024

An offline learning of behavior correction policy for vision-based robotic manipulation

ICRA 2024poster

Offline learning usually requires a large dataset for training. In this paper, we focus on vision-based robotic manipulation tasks and utilize certain task properties to achieve offline learning with a small dataset. We propose a two-stage agent consisting of a tentative decision stage and a correct…

Cited by 0SourceScholar
2024

Balancing Similarity and Complementarity for Federated Learning

ICML 2024poster

In mobile and IoT systems, Federated Learning (FL) is increasingly important for effectively using data while maintaining user privacy. One key challenge in FL is managing statistical heterogeneity, such as non-i.i.d. data, arising from numerous clients and diverse data sources. This requires strate…

Cited by 6SourcePDFScholar
2024

Counterfactual Reasoning for Multi-Label Image Classification via Patching-Based Training

ICML 2024poster

The key to multi-label image classification (MLC) is to improve model performance by leveraging label correlations. Unfortunately, it has been shown that overemphasizing co-occurrence relationships can cause the overfitting issue of the model, ultimately leading to performance degradation. In this p…

2024

Direct Distillation between Different Domains

ECCV 2024poster

"Knowledge Distillation (KD) aims to learn a compact student network using knowledge from a large pre-trained teacher network, where both networks are trained on data from the same distribution. However, in practical applications, the student network may be required to perform in a new scenario (i.e…

2024

Dual-Decoupling Learning and Metric-Adaptive Thresholding for Semi-Supervised Multi-Label Learning

ECCV 2024poster

"Semi-supervised multi-label learning (SSMLL) is a powerful framework for leveraging unlabeled data to reduce the expensive cost of collecting precise multi-label annotations. Unlike semi-supervised learning, one cannot select the most probable label as the pseudo-label in SSMLL due to multiple sema…

2024

Efficient Non-stationary Online Learning by Wavelets with Applications to Online Distribution Shift Adaptation

ICML 2024poster

Dynamic regret minimization offers a principled way for non-stationary online learning, where the algorithm's performance is evaluated against changing comparators. Prevailing methods often employ a two-layer online ensemble, consisting of a group of base learners with different configurations and a…

Cited by 2SourcePDFScholar
2024

Fixed-Budget Real-Valued Combinatorial Pure Exploration of Multi-Armed Bandit

AISTATS 2024poster

We study the real-valued combinatorial pure exploration of the multi-armed bandit in the fixed-budget setting. We first introduce an algorithm named the Combinatorial Successive Asign (CSA) algorithm, which is the first algorithm that can identify the best action even when the size of the action cla…

Cited by 1SourcePDFScholar
2024

Generating Chain-of-Thoughts with a Pairwise-Comparison Approach to Searching for the Most Promising Intermediate Thought

ICML 2024poster

To improve the ability of the large language model (LLMs) to tackle complex reasoning problems, chain-of-thoughts (CoT) methods were proposed to guide LLMs to reason step-by-step, enabling problem solving from simple to complex. State-of-the-art methods for generating such a chain involve interactiv…

Cited by 5SourcePDFScholar
2024

Imprecise Label Learning: A Unified Framework for Learning with Various Imprecise Label Configurations

NeurIPS 2024poster

Learning with reduced labeling standards, such as noisy label, partial label, and supplementary unlabeled data, which we generically refer to as imprecise label, is a commonplace challenge in machine learning tasks. Previous methods tend to propose specific designs for every emerging imprecise label…

2024

Learning with Complementary Labels Revisited: The Selected-Completely-at-Random Setting Is More Practical

ICML 2024poster

Complementary-label learning is a weakly supervised learning problem in which each training example is associated with one or multiple complementary labels indicating the classes to which it does not belong. Existing consistent approaches have relied on the uniform distribution assumption to model t…

2024

Locally Estimated Global Perturbations are Better than Local Perturbations for Federated Sharpness-aware Minimization

ICML 2024spotlight

In federated learning (FL), the multi-step update and data heterogeneity among clients often lead to a loss landscape with sharper minima, degenerating the performance of the resulted global model. Prevalent federated approaches incorporate sharpness-aware minimization (SAM) into local training to m…

2024

Robust Similarity Learning with Difference Alignment Regularization

ICLR 2024poster

Similarity-based representation learning has shown impressive capabilities in both supervised (e.g., metric learning) and unsupervised (e.g., contrastive learning) scenarios. Existing approaches effectively constrained the representation difference (i.e., the disagreement between the embeddings of t…

Cited by 0SourcePDFScholar
2024

Slight Corruption in Pre-training Data Makes Better Diffusion Models

NeurIPS 2024spotlight

Diffusion models (DMs) have shown remarkable capabilities in generating realistic high-quality images, audios, and videos. They benefit significantly from extensive pre-training on large-scale datasets, including web-crawled data with paired data and conditions, such as image-text and image-class p…

Cited by 6SourcePDFScholar
2024

Test-time Adaptation in Non-stationary Environments via Adaptive Representation Alignment

NeurIPS 2024poster

Adapting to distribution shifts is a critical challenge in modern machine learning, especially as data in many real-world applications accumulate continuously in the form of streams. We investigate the problem of sequentially adapting a model to non-stationary environments, where the data distributi…

Cited by 0SourcePDFScholar
2024

The Choice of Noninformative Priors for Thompson Sampling in Multiparameter Bandit Models

AAAI 2024technical

Thompson sampling (TS) has been known for its outstanding empirical performance supported by theoretical guarantees across various reward models in the classical stochastic multi-armed bandit problems. Nonetheless, its optimality is often restricted to specific priors due to the common observation t…

Cited by 0SourcePDFScholar
2024

Thompson Sampling for Real-Valued Combinatorial Pure Exploration of Multi-Armed Bandit

AAAI 2024technical

We study the real-valued combinatorial pure exploration of the multi-armed bandit (R-CPE-MAB) problem. In R-CPE-MAB, a player is given stochastic arms, and the reward of each arm follows an unknown distribution. In each time step, a player pulls a single arm and observes its reward. The player's goa…

Cited by 7SourcePDFScholar
2024

Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks

ICLR 2024spotlight

Pre-training on large-scale datasets and then fine-tuning on downstream tasks have become a standard practice in deep learning. However, pre-training data often contain label noise that may adversely affect the generalization of the model. This paper aims to understand the nature of noise in pre-tra…

2024

VEC-SBM: Optimal Community Detection with Vectorial Edges Covariates

AISTATS 2024poster

Social networks are often associated with rich side information, such as texts and images. While numerous methods have been developed to identify communities from pairwise interactions, they usually ignore such side information. In this work, we study an extension of the Stochastic Block Model (SBM)…

2024

Vision-Language Model Fine-Tuning via Simple Parameter-Efficient Modification

EMNLP 2024main

Recent advances in fine-tuning Vision-Language Models (VLMs) have witnessed the success of prompt tuning and adapter tuning, while the classic model fine-tuning on inherent parameters seems to be overlooked. It is believed that fine-tuning the parameters of VLMs with few-shot samples corrupts the pr…

2024

What Makes Partial-Label Learning Algorithms Effective?

NeurIPS 2024poster

A partial label (PL) specifies a set of candidate labels for an instance and partial-label learning (PLL) trains multi-class classifiers with PLs. Recently, many methods that incorporate techniques from other domains have shown strong potential. The expectation that stronger techniques would enhance…

Cited by 2SourcePDFScholar
2023

Adapting to Continuous Covariate Shift via Online Density Ratio Estimation

NeurIPS 2023poster

Dealing with distribution shifts is one of the central challenges for modern machine learning. One fundamental situation is the covariate shift, where the input distributions of data change from the training to testing stages while the input-conditional output distribution remains unchanged. In this…

Cited by 17SourcePDFScholar
2023

Binary Classification with Confidence Difference

NeurIPS 2023poster

Recently, learning with soft labels has been shown to achieve better performance than learning with hard labels in terms of model generalization, calibration, and robustness. However, collecting pointwise labeling confidence for all training examples can be challenging and time-consuming in real-wor…

Cited by 11SourcePDFScholar
2023

Class-Distribution-Aware Pseudo-Labeling for Semi-Supervised Multi-Label Learning

NeurIPS 2023poster

Pseudo-labeling has emerged as a popular and effective approach for utilizing unlabeled data. However, in the context of semi-supervised multi-label learning (SSMLL), conventional pseudo-labeling methods encounter difficulties when dealing with instances associated with multiple labels and an unknow…

2023

Distribution Shift Matters for Knowledge Distillation with Webly Collected Images

ICCV 2023poster

Knowledge distillation aims to learn a lightweight student network from a pre-trained teacher network. In practice, existing knowledge distillation methods are usually infeasible when the original training data is unavailable due to some privacy issues and data management considerations. Therefore,…

Cited by 17PDFScholar
2023

Distributional Pareto-Optimal Multi-Objective Reinforcement Learning

NeurIPS 2023poster

Multi-objective reinforcement learning (MORL) has been proposed to learn control policies over multiple competing objectives with each possible preference over returns. However, current MORL algorithms fail to account for distributional preferences over the multi-variate returns, which are particula…

2023

Diversified Outlier Exposure for Out-of-Distribution Detection via Informative Extrapolation

NeurIPS 2023poster

Out-of-distribution (OOD) detection is important for deploying reliable machine learning models on real-world applications. Recent advances in outlier exposure have shown promising results on OOD detection via fine-tuning model with informatively sampled auxiliary outliers. However, previous methods…

2023

Diversity-enhancing Generative Network for Few-shot Hypothesis Adaptation

ICML 2023poster

Generating unlabeled data has been recently shown to help address the few-shot hypothesis adaptation (FHA) problem, where we aim to train a classifier for the target domain with a few labeled target-domain data and a well-trained source-domain classifier (i.e., a source hypothesis), for the addition…

Cited by 4SourcePDFScholar
2023

Efficient Adversarial Contrastive Learning via Robustness-Aware Coreset Selection

NeurIPS 2023spotlight

Adversarial contrastive learning (ACL) does not require expensive data annotations but outputs a robust representation that withstands adversarial attacks and also generalizes to a wide range of downstream tasks. However, ACL needs tremendous running time to generate the adversarial variants of all…

2023

Enhancing Adversarial Contrastive Learning via Adversarial Invariant Regularization

NeurIPS 2023poster

Adversarial contrastive learning (ACL) is a technique that enhances standard contrastive learning (SCL) by incorporating adversarial data to learn a robust representation that can withstand adversarial attacks and common corruptions without requiring costly annotations. To improve transferability, t…

2023

GAT: Guided Adversarial Training with Pareto-optimal Auxiliary Tasks

ICML 2023poster

While leveraging additional training data is well established to improve adversarial robustness, it incurs the unavoidable cost of data collection and the heavy computation to train models. To mitigate the costs, we propose *Guided Adversarial Training * (GAT), a novel adversarial training technique…

2023

Generalizing Importance Weighting to A Universal Solver for Distribution Shift Problems

NeurIPS 2023spotlight

Distribution shift (DS) may have two levels: the distribution itself changes, and the support (i.e., the set where the probability density is non-zero) also changes. When considering the support change between the training and test distributions, there can be four cases: (i) they exactly match; (ii)…

2023

Is the Performance of My Deep Network Too Good to Be True? A Direct Approach to Estimating the Bayes Error in Binary Classification

ICLR 2023top-5%

There is a fundamental limitation in the prediction performance that a machine learning model can achieve due to the inevitable uncertainty of the prediction target. In classification problems, this can be characterized by the Bayes error, which is the best achievable error with any classifier. The…

2023

Multi-Label Knowledge Distillation

ICCV 2023poster

Existing knowledge distillation methods typically work by imparting the knowledge of output logits or intermediate feature maps from the teacher network to the student network, which is very successful in multi-class single-label learning. However, these methods can hardly be extended to the multi-l…

Cited by 24PDFcodeScholar
2023

On the Overlooked Pitfalls of Weight Decay and How to Mitigate Them: A Gradient-Norm Perspective

NeurIPS 2023poster

Weight decay is a simple yet powerful regularization technique that has been very widely used in training of deep neural networks (DNNs). While weight decay has attracted much attention, previous studies fail to discover some overlooked pitfalls on large gradient norms resulted by weight decay. In t…

2023

Online (Multinomial) Logistic Bandit: Improved Regret and Constant Computation Cost

NeurIPS 2023spotlight

This paper investigates the logistic bandit problem, a variant of the generalized linear bandit model that utilizes a logistic model to depict the feedback from an action. While most existing research focuses on the binary logistic bandit problem, the multinomial case, which considers more than two…

Cited by 18SourcePDFScholar
2023

Optimality of Thompson Sampling with Noninformative Priors for Pareto Bandits

ICML 2023poster

In the stochastic multi-armed bandit problem, a randomized probability matching policy called Thompson sampling (TS) has shown excellent performance in various reward models. In addition to the empirical performance, TS has been shown to achieve asymptotic problem-dependent lower bounds in several m…

Cited by 5SourcePDFScholar
2023

Seeing Differently, Acting Similarly: Heterogeneously Observable Imitation Learning

ICLR 2023top-25%

In many real-world imitation learning tasks, the demonstrator and the learner have to act under different observation spaces. This situation brings significant obstacles to existing imitation learning approaches, since most of them learn policies under homogeneous observation spaces. On the other ha…

Cited by 10SourcePDFScholar
2022

Adapting to Online Label Shift with Provable Guarantees

NeurIPS 2022accept

The standard supervised learning paradigm works effectively when training data shares the same distribution as the upcoming testing samples. However, this stationary assumption is often violated in real-world applications, especially when testing data appear in an online fashion. In this paper, we f…

Cited by 36SourcePDFScholar
2022

Adaptive Inertia: Disentangling the Effects of Adaptive Learning Rate and Momentum

ICML 2022oral

Adaptive Moment Estimation (Adam), which combines Adaptive Learning Rate and Momentum, would be the most popular stochastic optimizer for accelerating the training of deep neural networks. However, it is empirically known that Adam often generalizes worse than Stochastic Gradient Descent (SGD). The…

Cited by 69SourcePDFScholar
2022

Adversarial Attack and Defense for Non-Parametric Two-Sample Tests

ICML 2022spotlight

Non-parametric two-sample tests (TSTs) that judge whether two sets of samples are drawn from the same distribution, have been widely used in the analysis of critical data. People tend to employ TSTs as trusted basic tools and rarely have any doubt about their reliability. This paper systematically u…

2022

Adversarial Training with Complementary Labels: On the Benefit of Gradually Informative Attacks

NeurIPS 2022accept

Adversarial training (AT) with imperfect supervision is significant but receives limited attention. To push AT towards more practical scenarios, we explore a brand new yet challenging setting, i.e., AT with complementary labels (CLs), which specify a class that a data sample does not belong to. Howe…

2022

Exploiting Class Activation Value for Partial-Label Learning

ICLR 2022poster

Partial-label learning (PLL) solves the multi-class classification problem, where each training instance is assigned a set of candidate labels that include the true label. Recent advances showed that PLL can be compatible with deep neural networks, which achieved state-of-the-art performance. Howeve…

Cited by 59SourcePDFScholar
2022

Federated Learning from Only Unlabeled Data with Class-conditional-sharing Clients

ICLR 2022poster

Supervised federated learning (FL) enables multiple clients to share the trained model without sharing their labeled data. However, potential clients might even be reluctant to label their own data, which could limit the applicability of FL in practice. In this paper, we show the possibility of unsu…

2022

Instance-Dependent Label-Noise Learning With Manifold-Regularized Transition Matrix Estimation

CVPR 2022poster

In label-noise learning, estimating the transition matrix has attracted more and more attention as the matrix plays an important role in building statistically consistent classifiers. However, it is very challenging to estimate the transition matrix T(x), where T(x) denotes the instance, because it…

Cited by 93PDFScholar
2022

Learning Contrastive Embedding in Low-Dimensional Space

NeurIPS 2022accept

Contrastive learning (CL) pretrains feature embeddings to scatter instances in the feature space so that the training data can be well discriminated. Most existing CL techniques usually encourage learning such feature embeddings in the highdimensional space to maximize the instance discrimination. H…

Cited by 15SourcePDFScholar
2022

Meta Discovery: Learning to Discover Novel Classes given Very Limited Data

ICLR 2022spotlight

In novel class discovery (NCD), we are given labeled data from seen classes and unlabeled data from unseen classes, and we train clustering models for the unseen classes. However, the implicit assumptions behind NCD are still unclear. In this paper, we demystify assumptions behind NCD and find that…

2022

Pairwise Supervision Can Provably Elicit a Decision Boundary

AISTATS 2022poster

Similarity learning is a general problem to elicit useful representations by predicting the relationship between a pair of patterns. This problem is related to various important preprocessing tasks such as metric learning, kernel learning, and contrastive learning. A classifier built upon the repres…

Cited by 12SourcePDFScholar
2022

Predictive variational Bayesian inference as risk-seeking optimization

AISTATS 2022poster

Since the Bayesian inference works poorly under model misspecification, various solutions have been explored to counteract the shortcomings. Recently proposed predictive Bayes (PB) that directly optimizes the Kullback Leibler divergence between the empirical distribution and the approximate predicti…

Cited by 3SourcePDFScholar
2022

Rethinking Class-Prior Estimation for Positive-Unlabeled Learning

ICLR 2022poster

Given only positive (P) and unlabeled (U) data, PU learning can train a binary classifier without any negative data. It has two building blocks: PU class-prior estimation (CPE) and PU classification; the latter has been well studied while the former has received less attention. Hitherto, the distrib…

Cited by 26SourcePDFScholar
2022

Sample Selection with Uncertainty of Losses for Learning with Noisy Labels

ICLR 2022poster

In learning with noisy labels, the sample selection approach is very popular, which regards small-loss data as correctly labeled data during training. However, losses are generated on-the-fly based on the model being trained with noisy labels, and thus large-loss data are likely but not certain to be…

Cited by 159SourcePDFScholar
2022

Synergy-of-Experts: Collaborate to Improve Adversarial Robustness

NeurIPS 2022accept

Learning adversarially robust models require invariant predictions to a small neighborhood of its natural inputs, often encountering insufficient model capacity. There is research showing that learning multiple sub-models in an ensemble could mitigate this insufficiency, further improving the genera…

Cited by 8SourcePDFScholar
2022

To Smooth or Not? When Label Smoothing Meets Noisy Labels

ICML 2022oral

Label smoothing (LS) is an arising learning paradigm that uses the positively weighted average of both the hard training labels and uniformly distributed soft labels. It was shown that LS serves as a regularizer for training data with hard labels and therefore improves the generalization of the mode…

2022

Towards Adversarially Robust Deep Image Denoising

IJCAI 2022poster

This work systematically investigates the adversarial robustness of deep image denoisers (DIDs), i.e, how well DIDs can recover the ground truth from noisy observations degraded by adversarial perturbations. Firstly, to evaluate DIDs’ robustness, we propose a novel adversarial attack, namely Observa…

Cited by 16SourcePDFScholar
2021

A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima

ICLR 2021poster

Stochastic Gradient Descent (SGD) and its variants are mainstream methods for training deep networks in practice. SGD is known to find a flat minimum that often generalizes well. However, it is mathematically unclear how deep learning can select a flat minimum among so many minima. To answer the que…

Cited by 159SourcePDFScholar
2021

Binary Classification from Multiple Unlabeled Datasets via Surrogate Set Classification

ICML 2021spotlight

To cope with high annotation costs, training a classifier only from weakly supervised data has attracted a great deal of attention these days. Among various approaches, strengthening supervision from completely unsupervised classification is a promising direction, which typically employs class prior…

2021

CIFS: Improving Adversarial Robustness of CNNs via Channel-wise Importance-based Feature Selection

ICML 2021spotlight

We investigate the adversarial robustness of CNNs from the perspective of channel-wise activations. By comparing normally trained and adversarially trained models, we observe that adversarial training (AT) robustifies CNNs by aligning the channel-wise activations of adversarial data with those of th…

Cited by 56SourcePDFScholar
2021

Classification with Rejection Based on Cost-sensitive Classification

ICML 2021spotlight

The goal of classification with rejection is to avoid risky misclassification in error-critical applications such as medical diagnosis and product inspection. In this paper, based on the relationship between classification with rejection and cost-sensitive classification, we propose a novel method o…

Cited by 107SourcePDFScholar
2021

Confidence Scores Make Instance-dependent Label-noise Learning Possible

ICML 2021oral

In learning with noisy labels, for every instance, its label can randomly walk to other classes following a transition distribution which is named a noise model. Well-studied noise models are all instance-independent, namely, the transition depends only on the original label but not the instance its…

Cited by 137SourcePDFScholar
2021

Fenchel-Young Losses with Skewed Entropies for Class-posterior Probability Estimation

AISTATS 2021poster

We study class-posterior probability estimation (CPE) for binary responses where one class has much fewer data than the other. For example, events such as species co-occurrence in ecology and wars in political science are often much rarer than non-events. Logistic regression has been widely used for…

Cited by 10SourcePDFScholar
2021

Geometry-aware Instance-reweighted Adversarial Training

ICLR 2021oral

In adversarial machine learning, there was a common belief that robustness and accuracy hurt each other. The belief was challenged by recent studies where we can maintain the robustness and improve the accuracy. However, the other direction, whether we can keep the accuracy and improve the robustnes…

Cited by 339SourcePDFScholar
2021

Incorporating causal graphical prior knowledge into predictive modeling via simple data augmentation

UAI 2021poster

Causal graphs (CGs) are compact representations of the knowledge of the data generating processes behind the data distributions. When a CG is available, e.g., from the domain knowledge, we can infer the conditional independence (CI) relations that should hold in the data distribution. However, it is…

2021

Learning Diverse-Structured Networks for Adversarial Robustness

ICML 2021spotlight

In adversarial training (AT), the main focus has been the objective and optimizer while the model has been less studied, so that the models being used are still those classic ones in standard training (ST). Classic network architectures (NAs) are generally worse than searched NA in ST, which should…

2021

Learning Noise Transition Matrix from Only Noisy Labels via Total Variation Regularization

ICML 2021oral

Many weakly supervised classification methods employ a noise transition matrix to capture the class-conditional label corruption. To estimate the transition matrix from noisy data, existing methods often need to estimate the noisy class-posterior, which could be unreliable due to the overconfidence…

2021

Loss function based second-order Jensen inequality and its application to particle variational inference

NeurIPS 2021poster

Bayesian model averaging, obtained as the expectation of a likelihood function by a posterior distribution, has been widely used for prediction, evaluation of uncertainty, and model selection. Various approaches have been developed to efficiently capture the information in the posterior distribution…

Cited by 6SourcePDFScholar
2021

Lower-Bounded Proper Losses for Weakly Supervised Classification

ICML 2021spotlight

This paper discusses the problem of weakly supervised classification, in which instances are given weak labels that are produced by some label-corruption process. The goal is to derive conditions under which loss functions for weak-label learning are proper and lower-bounded—two essential requiremen…

2021

Maximum Mean Discrepancy Test is Aware of Adversarial Attacks

ICML 2021spotlight

The maximum mean discrepancy (MMD) test could in principle detect any distributional discrepancy between two datasets. However, it has been shown that the MMD test is unaware of adversarial attacks–the MMD test failed to detect the discrepancy between natural data and adversarial data. Given this ph…

2021

Mediated Uncoupled Learning: Learning Functions without Direct Input-output Correspondences

ICML 2021spotlight

Ordinary supervised learning is useful when we have paired training data of input $X$ and output $Y$. However, such paired data can be difficult to collect in practice. In this paper, we consider the task of predicting $Y$ from $X$ when we have no paired data of them, but we have two separate, indep…

2021

On Focal Loss for Class-Posterior Probability Estimation: A Theoretical Perspective

CVPR 2021poster

The focal loss has demonstrated its effectiveness in many real-world applications such as object detection and image classification, but its theoretical understanding has been limited so far. In this paper, we first prove that the focal loss is classification-calibrated, i.e., its minimizer surely y…

Cited by 34PDFScholar
2021

Pointwise Binary Classification with Pairwise Confidence Comparisons

ICML 2021spotlight

To alleviate the data requirement for training effective binary classifiers in binary classification, many weakly supervised learning settings have been proposed. Among them, some consider using pairwise but not pointwise labels, when pointwise labels are not accessible due to privacy, confidentiali…

Cited by 33SourcePDFScholar
2021

Positive-Negative Momentum: Manipulating Stochastic Gradient Noise to Improve Generalization

ICML 2021spotlight

It is well-known that stochastic gradient noise (SGN) acts as implicit regularization for deep learning and is essentially important for both optimization and generalization of deep networks. Some works attempted to artificially simulate SGN by injecting random noise to improve deep learning. Howeve…

2021

Probabilistic Margins for Instance Reweighting in Adversarial Training

NeurIPS 2021poster

Reweighting adversarial data during training has been recently shown to improve adversarial robustness, where data closer to the current decision boundaries are regarded as more critical and given larger weights. However, existing methods measuring the closeness are not very reliable: they are discr…

2021

Provably End-to-end Label-noise Learning without Anchor Points

ICML 2021spotlight

In label-noise learning, the transition matrix plays a key role in building statistically consistent classifiers. Existing consistent estimators for the transition matrix have been developed by exploiting anchor points. However, the anchor-point assumption is not always satisfied in real scenarios.…

2021

Robust Imitation Learning from Noisy Demonstrations

AISTATS 2021poster

Robust learning from noisy demonstrations is a practical but highly challenging problem in imitation learning. In this paper, we first theoretically show that robust imitation learning can be achieved by optimizing a classification risk with a symmetric loss. Based on this theoretical finding, we th…

2021

γ-ABC: Outlier-Robust Approximate Bayesian Computation Based on a Robust Divergence Estimator

AISTATS 2021poster

Approximate Bayesian computation (ABC) is a likelihood-free inference method that has been employed in various applications. However, ABC can be sensitive to outliers if a data discrepancy measure is chosen inappropriately. In this paper, we propose to use a nearest-neighbor-based γ-divergence estim…

Cited by 19SourcePDFScholar
2020

Accelerating the diffusion-based ensemble sampling by non-reversible dynamics

ICML 2020poster

Posterior distribution approximation is a central task in Bayesian inference. Stochastic gradient Langevin dynamics (SGLD) and its extensions have been practically used and theoretically studied. While SGLD updates a single particle at a time, ensemble methods that update multiple particles simultan…

Cited by 21SourcePDFScholar
2020

Analysis and Design of Thompson Sampling for Stochastic Partial Monitoring

NeurIPS 2020poster

We investigate finite stochastic partial monitoring, which is a general model for sequential learning with limited feedback. While Thompson sampling is one of the most promising algorithms on a variety of online decision-making problems, its properties for stochastic partial monitoring have not been…

Cited by 11SourcePDFScholar
2020

Attacks Which Do Not Kill Training Make Adversarial Learning Stronger

ICML 2020poster

Adversarial training based on the minimax formulation is necessary for obtaining adversarial robustness of trained models. However, it is conservative or even pessimistic so that it sometimes hurts the natural generalization. In this paper, we raise a fundamental question{—}do we have to trade off n…

Cited by 505SourcePDFScholar
2020

Binary Classification from Positive Data with Skewed Confidence

IJCAI 2020poster

Positive-confidence (Pconf) classification [Ishida et al., 2018] is a promising weakly-supervised learning method which trains a binary classifier only from positive data equipped with confidence. However, in practice, the confidence may be skewed by bias arising in an annotation process. The Pconf…

Cited by 0SourcePDFScholar
2020

Calibrated Surrogate Maximization of Linear-fractional Utility in Binary Classification

AISTATS 2020poster

Complex classification performance metrics such as the F-measure and Jaccard index are often used, in order to handle class-imbalanced cases such as information retrieval and image segmentation. These performance metrics are not decomposable, that is, they cannot be expressed in a per-example mann…

Cited by 19SourcePDFScholar
2020

Coupling-based Invertible Neural Networks Are Universal Diffeomorphism Approximators

NeurIPS 2020oral

Invertible neural networks based on coupling flows (CF-INNs) have various machine learning applications such as image synthesis and representation learning. However, their desirable characteristics such as analytic invertibility come at the cost of restricting the functional forms. This poses a ques…

Cited by 137SourcePDFScholar
2020

Do We Need Zero Training Loss After Achieving Zero Training Error?

ICML 2020poster

Overparameterized deep networks have the capacity to memorize training data with zero \emph{training error}. Even after memorization, the \emph{training loss} continues to approach zero, making the model overconfident and the test performance degraded. Since existing regularizers do not directly aim…

2020

Dual T: Reducing Estimation Error for Transition Matrix in Label-noise Learning

NeurIPS 2020poster

The transition matrix, denoting the transition relationship from clean labels to noisy labels, is essential to build statistically consistent classifiers in label-noise learning. Existing methods for estimating the transition matrix rely heavily on estimating the noisy class posterior. However, the…

Cited by 300SourcePDFScholar
2020

Learning from Aggregate Observations

NeurIPS 2020poster

We study the problem of learning from aggregate observations where supervision signals are given to sets of instances instead of individual instances, while the goal is still to predict labels of unseen individuals. A well-known example is multiple instance learning (MIL). In this paper, we extend…

2020

Mitigating Overfitting in Supervised Classification from Two Unlabeled Datasets: A Consistent Risk Correction Approach

AISTATS 2020poster

The recently proposed unlabeled-unlabeled (UU) classification method allows us to train a binary classifier only from two unlabeled datasets with different class priors. Since this method is based on the empirical risk minimization, it works as if it is a supervised classification method, compatible…

Cited by 71SourcePDFScholar
2020

Normalized Flat Minima: Exploring Scale Invariant Definition of Flat Minima for Neural Networks Using PAC-Bayesian Analysis

ICML 2020poster

The notion of flat minima has gained attention as a key metric of the generalization ability of deep learning models. However, current definitions of flatness are known to be sensitive to parameter rescaling. While some previous studies have proposed to rescale flatness metrics using parameter scale…

Cited by 81SourcePDFScholar
2020

Online Dense Subgraph Discovery via Blurred-Graph Feedback

ICML 2020poster

Dense subgraph discovery aims to find a dense component in edge-weighted graphs. This is a fundamental graph-mining task with a variety of applications and thus has received much attention recently. Although most existing methods assume that each individual edge weight is easily obtained, such an as…

Cited by 17SourcePDFScholar
2020

Part-dependent Label Noise: Towards Instance-dependent Label Noise

NeurIPS 2020spotlight

Learning with the \textit{instance-dependent} label noise is challenging, because it is hard to model such real-world noise. Note that there are psychological and physiological evidences showing that we humans perceive instances by decomposing them into parts. Annotators are therefore more likely to…

2020

Progressive Identification of True Labels for Partial-Label Learning

ICML 2020poster

Partial-label learning (PLL) is a typical weakly supervised learning problem, where each training instance is equipped with a set of candidate labels among which only one is the true label. Most existing methods elaborately designed learning objectives as constrained optimizations that must be solve…

2020

Rethinking Importance Weighting for Deep Learning under Distribution Shift

NeurIPS 2020spotlight

Under distribution shift (DS) where the training data distribution differs from the test one, a powerful technique is importance weighting (IW) which handles DS in two separate steps: weight estimation (WE) estimates the test-over-training density ratio and weighted classification (WC) trains the cl…

2020

SIGUA: Forgetting May Make Learning with Noisy Labels More Robust

ICML 2020poster

Given data with noisy labels, over-parameterized deep networks can gradually memorize the data, and fit everything in the end. Although equipped with corrections for noisy labels, many learning methods in this area still suffer overfitting due to undesired memorization. In this paper, to relieve thi…

Cited by 160SourcePDFScholar
2020

Simultaneous Planning for Item Picking and Placing by Deep Reinforcement Learning

IROS 2020poster

Container loading by a picking robot is an important challenge in the logistics industry. When designing such a robotic system, item picking and placing have been planned individually thus far. However, since the condition of picking an item affects the possible candidates for placing, it is prefera…

Cited by 14SourceScholar
2020

Unbiased Risk Estimators Can Mislead: A Case Study of Learning with Complementary Labels

ICML 2020poster

In weakly supervised learning, unbiased risk estimator(URE) is a powerful tool for training classifiers when training and test data are drawn from different distributions. Nevertheless, UREs lead to overfitting in many problem settings when the models are complex like deep networks. In this paper, w…

Cited by 69SourcePDFScholar
2020

Variational Imitation Learning with Diverse-quality Demonstrations

ICML 2020poster

Learning from demonstrations can be challenging when the quality of demonstrations is diverse, and even more so when the quality is unknown and there is no additional information to estimate the quality. We propose a new method for imitation learning in such scenarios. We show that simple quality-es…

2019

Are Anchor Points Really Indispensable in Label-Noise Learning?

NeurIPS 2019poster

In label-noise learning, the \textit{noise transition matrix}, denoting the probabilities that clean labels flip into noisy labels, plays a central role in building \textit{statistically consistent classifiers}. Existing theories have shown that the transition matrix can be learned by exploiting \te…

2019

Binary Classification Only from Unlabeled Data by Iterative Unlabeled-unlabeled Classification

ICASSP 2019accepted

Unlabeled-unlabeled (UU) classification (du Plessis et al. 2013) allows us to train a binary classifier from two sets of unlabeled data with different class priors. In this paper, we go beyond this scenario and try to train a binary classifier only from a single set of unlabeled data. Our key idea i…

Cited by 0SourceScholar
2019

Complementary-Label Learning for Arbitrary Losses and Models

ICML 2019oral

In contrast to the standard classification paradigm where the true class is given to each training pattern, complementary-label learning only uses training patterns each equipped with a complementary label, which only specifies one of the classes that the pattern does not belong to. The goal of this…

2019

Hierarchical Reinforcement Learning via Advantage-Weighted Information Maximization

ICLR 2019poster

Real-world tasks are often highly structured. Hierarchical reinforcement learning (HRL) has attracted research interest as an approach for leveraging the hierarchical structure of a given task in reinforcement learning (RL). However, identifying the hierarchical policy structure that enhances the pe…

2019

How does Disagreement Help Generalization against Label Corruption?

ICML 2019oral

Learning with noisy labels is one of the hottest problems in weakly-supervised learning. Based on memorization effects of deep neural networks, training on small-loss instances becomes very promising for handling noisy labels. This fosters the state-of-the-art approach "Co-teaching" that cross-train…

Cited by 975SourcePDFScholar
2019

Imitation Learning from Imperfect Demonstration

ICML 2019oral

Imitation learning (IL) aims to learn an optimal policy from demonstrations. However, such demonstrations are often imperfect since collecting optimal ones is costly. To effectively learn from imperfect demonstrations, we propose a novel approach that utilizes confidence scores, which describe the q…

Cited by 201SourcePDFScholar
2019

Learning Efficient Tensor Representations with Ring-structured Networks

ICASSP 2019accepted

Tensor train decomposition is a powerful representation for high-order tensors, which has been successfully applied to various machine learning tasks in recent years. In this paper, we study a more generalized tensor decomposition with a ring-structured network by employing circular multilinear prod…

Cited by 0SourceScholar
2019

On the Calibration of Multiclass Classification with Rejection

NeurIPS 2019poster

We investigate the problem of multiclass classification with rejection, where a classifier can choose not to make a prediction to avoid critical misclassification. First, we consider an approach based on simultaneous training of a classifier and a rejector, which achieves the state-of-the-art perfor…

2019

On the Minimal Supervision for Training Any Binary Classifier from Only Unlabeled Data

ICLR 2019poster

Empirical risk minimization (ERM), with proper loss function and regularization, is the common practice of supervised classification. In this paper, we study training arbitrary (from linear to deep) binary classifier from only unlabeled (U) data by ERM. We prove that it is impossible to estimate the…

2019

Uncoupled Regression from Pairwise Comparison Data

NeurIPS 2019poster

Uncoupled regression is the problem to learn a model from unlabeled data and the set of target values while the correspondence between them is unknown. Such a situation arises in predicting anonymized targets that involve sensitive information, e.g., one's annual income. Since existing methods for u…

2018

A fully adaptive algorithm for pure exploration in linear bandits

AISTATS 2018poster

We propose the first fully-adaptive algorithm for pure exploration in linear bandits—the task to find the arm with the largest expected reward, which depends on an unknown parameter linearly. While existing methods partially or entirely fix sequences of arm selections before observing rewards, our…

2018

Analysis of Minimax Error Rate for Crowdsourcing and Its Application to Worker Clustering Model

ICML 2018oral

While crowdsourcing has become an important means to label data, there is great interest in estimating the ground truth from unreliable labels produced by crowdworkers. The Dawid and Skene (DS) model is one of the most well-known models in the study of crowdsourcing. Despite its practical popularity…

2018

Bayesian Nonparametric Poisson-Process Allocation for Time-Sequence Modeling

AISTATS 2018poster

Analyzing the underlying structure of multiple time-sequences provides insights into the understanding of social networks and human activities. In this work, we present the Bayesian nonparametric Poisson process allocation (BaNPPA), a latent-function model for time-sequences, which automatically in…

2018

Co-teaching: Robust training of deep neural networks with extremely noisy labels

NeurIPS 2018poster

Deep learning with noisy labels is practically challenging, as the capacity of deep models is so high that they can totally memorize these noisy labels sooner or later during training. Nonetheless, recent studies on the memorization effects of deep neural networks show that they would first memorize…

2018

Continuous-time Value Function Approximation in Reproducing Kernel Hilbert Spaces

NeurIPS 2018poster

Motivated by the success of reinforcement learning (RL) for discrete-time tasks such as AlphaGo and Atari games, there has been a recent surge of interest in using RL for continuous-time control of physical systems (cf. many challenging tasks in OpenAI Gym and DeepMind Control Suite). Since discreti…

2018

Does Distributionally Robust Supervised Learning Give Robust Classifiers?

ICML 2018oral

Distributionally Robust Supervised Learning (DRSL) is necessary for building reliable machine learning systems. When machine learning is deployed in the real world, its performance can be significantly degraded because test data may follow a different distribution from training data. DRSL with f-div…

Cited by 347SourcePDFScholar
2018

Lipschitz-Margin Training: Scalable Certification of Perturbation Invariance for Deep Neural Networks

NeurIPS 2018poster

High sensitivity of neural networks against malicious perturbations on inputs causes security concerns. To take a steady step towards robust classifiers, we aim to create neural network models provably defended from perturbations. Prior certification work requires strong assumptions on network struc…

2018

Masking: A New Perspective of Noisy Supervision

NeurIPS 2018poster

It is important to learn various types of classifiers given training data with noisy labels. Noisy labels, in the most popular noise model hitherto, are corrupted from ground-truth labels by an unknown noise transition matrix. Thus, by estimating this matrix, classifiers can escape from overfitting…

2018

Multi Task Learning with Positive and Unlabeled Data and its Application to Mental State Prediction

ICASSP 2018accepted

In real-world machine learning applications, we are often faced with a situation where only a small number of training samples is available due to high sampling costs. For instance, prediction of mental states such as drowsiness from physiological information is a typical example. To cope with this…

Cited by 0SourceScholar
2017

Estimating Density Ridges by Direct Estimation of Density-Derivative-Ratios

AISTATS 2017poster

Estimation of \emphdensity ridges has been gathering a great deal of attention since it enables us to reveal lower-dimensional structures hidden in data. Recently, \emphsubspace constrained mean shift (SCMS) was proposed as a practical algorithm for density ridge estimation. A key technical ingredie…

Cited by 8SourcePDFScholar
2017

Generative Local Metric Learning for Kernel Regression

NeurIPS 2017poster

This paper shows how metric learning can be used with Nadaraya-Watson (NW) kernel regression. Compared with standard approaches, such as bandwidth selection, we show how metric learning can significantly reduce the mean square error (MSE) in kernel regression, particularly for high-dimensional data…

Cited by 20SourcePDFScholar
2017

Learning Discrete Representations via Information Maximizing Self-Augmented Training

ICML 2017poster

Learning discrete representations of data is a central machine learning task because of the compactness of the representations and ease of interpretation. The task includes clustering and hash learning as special cases. Deep neural networks are promising to be used because they can model the non-lin…

2017

Least-Squares Log-Density Gradient Clustering for Riemannian Manifolds

AISTATS 2017poster

Mean shift is a mode-seeking clustering algorithm that has been successfully used in a wide range of applications such as image segmentation and object tracking. To further improve the clustering performance, mean shift has been extended to various directions, including generalization to handle data…

Cited by 8SourcePDFScholar
2017

Positive-Unlabeled Learning with Non-Negative Risk Estimator

NeurIPS 2017oral

From only positive (P) and unlabeled (U) data, a binary classifier could be trained with PU learning, in which the state of the art is unbiased PU learning. However, if its model is very flexible, empirical risks on training data will go negative, and we will suffer from serious overfitting. In this…

2017

Semi-Supervised Classification Based on Classification from Positive and Unlabeled Data

ICML 2017poster

Most of the semi-supervised classification methods developed so far use unlabeled data for regularization purposes under particular distributional assumptions such as the cluster assumption. In contrast, recently developed methods of classification from positive and unlabeled data (PU classification…

Cited by 138SourcePDFScholar
2016

Non-Gaussian Component Analysis with Log-Density Gradient Estimation

AISTATS 2016poster

Non-Gaussian component analysis (NGCA) is aimed at identifying a linear subspace such that the projected data follows a non-Gaussian distribution. In this paper, we propose a novel NGCA algorithm based on log-density gradient estimation. Unlike existing methods, the proposed NGCA algorithm identifie…

Cited by 21SourcePDFScholar
2016

Theoretical Comparisons of Positive-Unlabeled Learning against Positive-Negative Learning

NeurIPS 2016poster

In PU learning, a binary classifier is trained from positive (P) and unlabeled (U) data without negative (N) data. Although N data is missing, it sometimes outperforms PN learning (i.e., ordinary supervised learning). Hitherto, neither theoretical nor experimental analysis has been given to explain…

Cited by 152SourcePDFScholar
2015

A dependence maximization approach towards street map-based localization

IROS 2015poster

In this paper, we present a novel approach to 2D street map-based localization for mobile robots that navigate mainly in urban sidewalk environments. Recently, localization based on the map built by Simultaneous Localization and Mapping (SLAM) has been widely used with great success. However, such m…

Cited by 11SourceScholar
2015

Convex Formulation for Learning from Positive and Unlabeled Data

ICML 2015poster

We discuss binary classification from only from positive and unlabeled data (PU classification), which is conceivable in various real-world machine learning problems. Since unlabeled data consists of both positive and negative data, simply separating positive and unlabeled data yields a biased solut…

Cited by 407SourcePDFScholar
2015

Direct Density-Derivative Estimation and Its Application in KL-Divergence Approximation

AISTATS 2015poster

Estimation of density derivatives is a versatile tool in statistical data analysis. A naive approach is to first estimate the density and then compute its derivative. However, such a two-step approach does not work well because a good density estimator does not necessarily mean a good density-deriv…

Cited by 28SourcePDFScholar