← Search

Shahana Ibrahim

10 accepted papers

2025

Multi-label Recognition under Noisy Supervision: A Confusion Mixture Modeling Approach

ICASSP 2025accepted

Multi-label recognition is a critical task in artificial intelligence, aiming to identify every object present in an image. Designing a multi-label classifier is a nontrivial task both from data collection and modeling perspectives. Collecting multiple labels for each image is extremely time-consumi…

Cited by 0SourceScholar
2025

Under-Counted Matrix Completion Without Detection Features

ICASSP 2025accepted

Under-counted matrix completion (UC-MC) has many important applications, especially in epidemiology and ecology where the observed data are often smaller than the actual numbers. Existing works model the under-counting effects using entry-wise miss detection probabilities, which are usually formulat…

Cited by 0SourceScholar
2024

Noisy Label Learning with Instance-Dependent Outliers: Identifiability via Crowd Wisdom

NeurIPS 2024spotlight

The generation of label noise is often modeled as a process involving a probability transition matrix (also interpreted as the _annotator confusion matrix_) imposed onto the label distribution. Under this model, learning the ``ground-truth classifier''---i.e., the classifier that can be learned if n…

Cited by 1SourcePDFScholar
2023

Deep Clustering with Incomplete Noisy Pairwise Annotations: A Geometric Regularization Approach

ICML 2023poster

The recent integration of deep learning and pairwise similarity annotation-based constrained clustering---i.e., deep constrained clustering (DCC)---has proven effective for incorporating weak supervision into massive data clustering: Less than 1% of pair similarity annotations can often substantiall…

2023

Deep Learning From Crowdsourced Labels: Coupled Cross-Entropy Minimization, Identifiability, and Regularization

ICLR 2023poster

Using noisy crowdsourced labels from multiple annotators, a deep learning-based end-to-end (E2E) system aims to learn the label correction mechanism and the neural classifier simultaneously. To this end, many E2E systems concatenate the neural classifier with multiple annotator-specific label confus…

2023

Under-Counted Tensor Completion with Neural Incorporation of Attributes

ICML 2023poster

Systematic under-counting effects are observed in data collected across many disciplines, e.g., epidemiology and ecology. Under-counted tensor completion (UC-TC) is well-motivated for many data analytics tasks, e.g., inferring the case numbers of infectious diseases at unobserved locations from unde…

2021

Crowdsourcing via Annotator Co-occurrence Imputation and Provable Symmetric Nonnegative Matrix Factorization

ICML 2021oral

Unsupervised learning of the Dawid-Skene (D&S) model from noisy, incomplete and crowdsourced annotations has been a long-standing challenge, and is a critical step towards reliably labeling massive data. A recent work takes a coupled nonnegative matrix factorization (CNMF) perspective, and shows app…

Cited by 15SourcePDFScholar
2021

Fiber-Sampled Stochastic Mirror Descent for Tensor Decomposition with β-Divergence

ICASSP 2021accepted

Canonical polyadic decomposition (CPD) has been a workhorse for multimodal data analytics. This work puts forth a stochastic algorithmic framework for CPD under β-divergence, which is well-motivated in statistical learning—where the Euclidean distance is typically not preferred. Despite the existenc…

Cited by 0SourceScholar
2021

Learning Mixed Membership from Adjacency Graph Via Systematic Edge Query: Identifiability and Algorithm

ICASSP 2021accepted

Graph clustering is a core technique for network analysis problems, e.g., community detection. This work puts forth a node clustering approach for largely incomplete adjacency graphs. Under the considered scenario, instead of having access to the complete graph, only a small amount of queries about…

Cited by 0SourceScholar
2019

Crowdsourcing via Pairwise Co-occurrences: Identifiability and Algorithms

NeurIPS 2019poster

The data deluge comes with high demands for data labeling. Crowdsourcing (or, more generally, ensemble learning) techniques aim to produce accurate labels via integrating noisy, non-expert labeling from annotators. The classic Dawid-Skene estimator and its accompanying expectation maximization (EM)…

Cited by 44SourcePDFScholar