← Search

Ilya Shpitser

24 accepted papers

2025

Feature Importance Metrics in the Presence of Missing Data

ICML 2025poster

Feature importance metrics are critical for interpreting machine learning models and understanding the relevance of individual features. However, real-world data often exhibit missingness, thereby complicating how feature importance should be evaluated. We introduce the distinction between two eval…

Cited by 0SourcePDFScholar
2024

A General Identification Algorithm For Data Fusion Problems Under Systematic Selection

UAI 2024poster

Causal inference is made challenging by confounding, selection bias, and other complications. A common approach to addressing these difficulties is the inclusion of auxiliary data on the superpopulation of interest. Such data may measure a different set of variables, or be obtained under different…

Cited by 2SourcePDFScholar
2024

Identification and Estimation for Nonignorable Missing Data: A Data Fusion Approach

ICML 2024poster

We consider the task of identifying and estimating a parameter of interest in settings where data is missing not at random (MNAR). In general, such parameters are not identified without strong assumptions on the missing data model. In this paper, we take an alternative approach and introduce a metho…

Cited by 1SourcePDFScholar
2024

Zero Inflation as a Missing Data Problem: a Proxy-based Approach

UAI 2024poster

A common type of zero-inflated data has certain true values incorrectly replaced by zeros due to data recording conventions (rare outcomes assumed to be absent) or details of data recording equipment (e.g. artificial zeros in gene expression data). Existing methods for zero-inflated data either fit…

2022

Causal Discovery in Linear Latent Variable Models Subject to Measurement Error

NeurIPS 2022accept

We focus on causal discovery in the presence of measurement error in linear systems where the mixing matrix, i.e., the matrix indicating the independent exogenous noise terms pertaining to the observed variables, is identified up to permutation and scaling of the columns. We demonstrate a somewhat s…

2022

Minimax Kernel Machine Learning for a Class of Doubly Robust Functionals with Application to Proximal Causal Inference

AISTATS 2022poster

Robins et al. (2008) introduced a class of influence functions (IFs) which could be used to obtain doubly robust moment functions for the corresponding parameters. However, that class does not include the IF of parameters for which the nuisance functions are solutions to integral equations. Such par…

2022

Semiparametric causal sufficient dimension reduction of multidimensional treatments

UAI 2022poster

Cause-effect relationships are typically evaluated by comparing outcome responses to binary treatment values, representing two arms of a hypothetical randomized controlled trial. However, in certain applications, treatments of interest are continuous and multidimensional. For example, understanding…

2021

Differentiable Causal Discovery Under Unmeasured Confounding

AISTATS 2021poster

The data drawn from biological, economic, and social systems are often confounded due to the presence of unmeasured variables. Prior work in causal discovery has focused on discrete search procedures for selecting acyclic directed mixed graphs (ADMGs), specifically ancestral ADMGs, that encode ordin…

2021

Entropic Inequality Constraints from e-separation Relations in Directed Acyclic Graphs with Hidden Variables

UAI 2021poster

Directed acyclic graphs (DAGs) with hidde variables are often used to characterize causal relations between variables in a system. When some variables are unobserved, DAGs imply a notoriously complicated set of constraints on the distribution of observed variables. In this work, we present entropic…

Cited by 7SourcePDFScholar
2021

Partial Identifiability in Discrete Data with Measurement Error

UAI 2021poster

When data contains measurement errors, it is necessary to make modeling assumptions relating the error-prone measurements to the unobserved true values. Work on measurement error has largely focused on models that fully identify the parameter of interest. As a result, many practically useful models…

Cited by 9SourcePDFScholar
2020

Deriving Bounds And Inequality Constraints Using Logical Relations Among Counterfactuals

UAI 2020poster

Causal parameters may not be point identified in the presence of unobserved confounding. However, information about non-identified parameters, in the form of bounds, may still be recovered from the observed data in some cases. We develop a new general method for obtaining bounds on causal paramet…

Cited by 18SourcePDFScholar
2020

Full Law Identification in Graphical Models of Missing Data: Completeness Results

ICML 2020poster

Missing data has the potential to affect analyses conducted in all fields of scientific study including healthcare, economics, and the social sciences. Several approaches to unbiased inference in the presence of non-ignorable missingness rely on the specification of the target distribution and its m…

Cited by 67SourcePDFScholar
2020

General Identification of Dynamic Treatment Regimes Under Interference

AISTATS 2020poster

In many applied fields, researchers are ofteninterested in tailoring treatments to unit-levelcharacteristics in order to optimize an outcomeof interest. Methods for identifying andestimating treatment policies are the subjectof the dynamic treatment regime literature. Separately, in many settings th…

Cited by 13SourcePDFScholar
2020

Identification and Estimation of Causal Effects Defined by Shift Interventions

UAI 2020poster

Causal inference quantifies cause effect relationships by means of counterfactual responses had some variable been artificially set to a constant. A more refined notion of manipulation, where a variable is artificially set to a fixed function of its natural value is also of interest in particular do…

Cited by 13SourcePDFScholar
2019

A Potential Outcomes Calculus for Identifying Conditional Path-Specific Effects

AISTATS 2019poster

The do-calculus is a well-known deductive system for deriving connections between interventional and observed distributions, and has been proven complete for a number of important identifiability problems in causal inference. Nevertheless, as it is currently defined, the do-calculus is inapplicable…

Cited by 79SourcePDFScholar
2019

Identification In Missing Data Models Represented By Directed Acyclic Graphs

UAI 2019poster

Missing data is a pervasive problem in data analyses, resulting in datasets that contain censored realizations of a target distribution. Many approaches to inference on the target distribution using censored observed data, rely on missing data models represented as a factorization with respect to a…

Cited by 54SourcePDFScholar