← Search

Krzysztof Dembczynski

9 accepted papers

2025

Optimal downsampling for Imbalanced Classification with Generalized Linear Models

AISTATS 2025poster

Downsampling or under-sampling is a technique that is utilized in the context of large and highly imbalanced classification models. We study optimal downsampling for imbalanced classification using generalized linear models (GLMs). We propose a pseudo maximum likelihood estimator and study its asymp…

Cited by 0SourceScholar
2024

A General Online Algorithm for Optimizing Complex Performance Metrics

ICML 2024poster

We consider sequential maximization of performance metrics that are general functions of a confusion matrix of a classifier (such as precision, F-measure, or G-mean). Such metrics are, in general, non-decomposable over individual instances, making their optimization very challenging. While they have…

Cited by 0SourcePDFScholar
2024

Consistent algorithms for multi-label classification with macro-at-$k$ metrics

ICLR 2024poster

We consider the optimization of complex performance metrics in multi-label classification under the population utility framework. We mainly focus on metrics linearly decomposable into a sum of binary classification utilities applied separately to each label with an additional requirement of exactly…

2023

Generalized test utilities for long-tail performance in extreme multi-label classification

NeurIPS 2023poster

Extreme multi-label classification (XMLC) is the task of selecting a small subset of relevant labels from a very large set of possible labels. As such, it is characterized by long-tail labels, i.e., most labels have very few positive instances. With standard performance measures such as precision@k…

2022

Regret Bounds for Multilabel Classification in Sparse Label Regimes

NeurIPS 2022accept

Multi-label classification (MLC) has wide practical importance, but the theoretical understanding of its statistical properties is still limited. As an attempt to fill this gap, we thoroughly study upper and lower regret bounds for two canonical MLC performance measures, Hamming loss and Precision@$…

Cited by 2SourcePDFScholar
2021

Online probabilistic label trees

AISTATS 2021poster

We introduce online probabilistic label trees (OPLTs), an algorithm that trains a label tree classifier in a fully online manner without any prior knowledge about the number of training instances, their features and labels. OPLTs are characterized by low time and space complexity as well as strong t…

2018

A no-regret generalization of hierarchical softmax to extreme multi-label classification

NeurIPS 2018poster

Extreme multi-label classification (XMLC) is a problem of tagging an instance with a small subset of relevant labels chosen from an extremely large pool of possible labels. Large label spaces can be efficiently handled by organizing labels as a tree, like in the hierarchical softmax (HSM) approach c…

2016

Extreme F-measure Maximization using Sparse Probability Estimates

ICML 2016poster

We consider the problem of (macro) F-measure maximization in the context of extreme multi-label classification (XMLC), i.e., multi-label classification with extremely large label spaces. We investigate several approaches based on recent results on the maximization of complex performance measures in…

2015

Online F-Measure Optimization

NeurIPS 2015poster

The F-measure is an important and commonly used performance metric for binary prediction tasks. By combining precision and recall into a single score, it avoids disadvantages of simple metrics like the error rate, especially in cases of imbalanced class distributions. The problem of optimizing the F…

Cited by 49SourcePDFScholar