← Search

Pradeep Kumar Ravikumar

34 accepted papers

2025

Contextures: Representations from Contexts

ICML 2025poster

Despite the empirical success of foundation models, we do not have a systematic characterization of the representations that these models learn. In this paper, we establish the contexture theory. It shows that a large class of representation learning methods can be characterized as learning from th…

Cited by 0SourcePDFScholar
2025

Learning from weak labelers as constraints

ICLR 2025poster

We study programmatic weak supervision, where in contrast to labeled data, we have access to \emph{weak labelers}, each of which either abstains or provides noisy labels corresponding to any input. Most previous approaches typically employ latent generative models that model the joint distribution o…

Cited by 0SourcePDFScholar
2025

On the Consistent Recovery of Joint Distributions from Conditionals

AISTATS 2025poster

Self-supervised learning methods that mask parts of the input data and train models to predict the missing components have led to significant advances in machine learning. These approaches learn conditional distributions $p(x_T \mid x_S)$ simultaneously, where $x_S$ and $x_T$ are subsets of the obse…

Cited by 0SourceScholar
2024

Do LLMs dream of elephants (when told not to)? Latent concept association and associative memory in transformers

NeurIPS 2024poster

Large Language Models (LLMs) have the capacity to store and recall facts. Through experimentation with open-source models, we observe that this ability to retrieve facts can be easily manipulated by changing contexts, even without altering their factual meanings. These findings highlight that LLMs m…

Cited by 6SourcePDFScholar
2024

From Causal to Concept-Based Representation Learning

NeurIPS 2024poster

To build intelligent machine learning systems, modern representation learning attempts to recover latent generative factors from data, such as in causal representation learning. A key question in this growing field is to provide rigorous conditions under which latent factors can be identified and th…

Cited by 2SourcePDFScholar
2024

Identifying General Mechanism Shifts in Linear Causal Representations

NeurIPS 2024poster

We consider the linear causal representation learning setting where we observe a linear mixing of $d$ unknown latent factors, which follow a linear structural causal model. Recent work has shown that it is possible to recover the latent factors as well as the underlying structural causal model over…

2024

Identifying Representations for Intervention Extrapolation

ICLR 2024poster

The premise of identifiable and causal representation learning is to improve the current representation learning paradigm in terms of generalizability or robustness. Despite recent progress in questions of identifiability, more theoretical results demonstrating concrete advantages of these methods f…

Cited by 18SourcePDFScholar
2024

LogiCity: Advancing Neuro-Symbolic AI with Abstract Urban Simulation

NeurIPS 2024poster

Recent years have witnessed the rapid development of Neuro-Symbolic (NeSy) AI systems, which integrate symbolic reasoning into deep neural networks. However, most of the existing benchmarks for NeSy AI fail to provide long-horizon reasoning tasks with complex multi-agent interactions. Furthermore, t…

2024

Markov Equivalence and Consistency in Differentiable Structure Learning

NeurIPS 2024poster

Existing approaches to differentiable structure learning of directed acyclic graphs (DAGs) rely on strong identifiability assumptions in order to guarantee that global minimizers of the acyclicity-constrained optimization problem identifies the true DAG. Moreover, it has been observed empirically th…

2024

On the Origins of Linear Representations in Large Language Models

ICML 2024poster

An array of recent works have argued that high-level semantic concepts are encoded "linearly" in the representation space of large language models. In this work, we study the origins of such linear representations. To that end, we introduce a latent variable model to abstract and formalize the conce…

Cited by 25SourcePDFScholar
2024

Spectrally Transformed Kernel Regression

ICLR 2024spotlight

Unlabeled data is a key component of modern machine learning. In general, the role of unlabeled data is to impose a form of smoothness, usually from the similarity information encoded in a base kernel, such as the ϵ-neighbor kernel or the adjacency matrix of a graph. This work revisits the classical…

Cited by 3SourcePDFScholar
2024

Understanding Augmentation-based Self-Supervised Representation Learning via RKHS Approximation and Regression

ICLR 2024spotlight

Data augmentation is critical to the empirical success of modern self-supervised representation learning, such as contrastive learning and masked language modeling. However, a theoretical understanding of the exact role of the augmentation remains limited. Recent work has built the connection betwee…

Cited by 16SourcePDFScholar
2023

Concept Gradient: Concept-based Interpretation Without Linear Assumption

ICLR 2023poster

Concept-based interpretations of black-box models are often more intuitive for humans to understand. The most widely adopted approach for concept-based, gradient interpretation is Concept Activation Vector (CAV). CAV relies on learning a linear relation between some latent representation of a given…

2023

Global Optimality in Bivariate Gradient-based DAG Learning

NeurIPS 2023poster

Recently, a new class of non-convex optimization problems motivated by the statistical problem of learning an acyclic directed graphical model from data has attracted significant interest. While existing work uses standard first-order optimization schemes to solve this problem, proving the global op…

Cited by 9SourcePDFScholar
2023

Learning Linear Causal Representations from Interventions under General Nonlinear Mixing

NeurIPS 2023oral

We study the problem of learning causal representations from unknown, latent interventions in a general setting, where the latent distribution is Gaussian but the mixing function is completely general. We prove strong identifiability results given unknown single-node interventions, i.e., without hav…

Cited by 71SourcePDFScholar
2023

Learning with Explanation Constraints

NeurIPS 2023poster

As larger deep learning models are hard to interpret, there has been a recent focus on generating explanations of these black-box models. In contrast, we may have apriori explanations of how models should behave. In this paper, we formalize this notion as learning from explanation constraints and p…

Cited by 7SourcePDFScholar
2023

Optimizing NOTEARS Objectives via Topological Swaps

ICML 2023poster

Recently, an intriguing class of non-convex optimization problems has emerged in the context of learning directed acyclic graphs (DAGs). These problems involve minimizing a given loss or score function, subject to a non-convex continuous constraint that penalizes the presence of cycles in a graph. I…

2023

Representer Point Selection for Explaining Regularized High-dimensional Models

ICML 2023poster

We introduce a novel class of sample-based explanations we term *high-dimensional representers*, that can be used to explain the predictions of a regularized high-dimensional model in terms of importance weights for each of the training samples. Our workhorse is a novel representer theorem for gener…

Cited by 4SourcePDFScholar
2023

Responsible AI (RAI) Games and Ensembles

NeurIPS 2023poster

Several recent works have studied the societal effects of AI; these include issues such as fairness, robustness, and safety. In many of these objectives, a learner seeks to minimize its worst-case loss over a set of predefined distributions (known as uncertainty sets), with usual examples being per…

2023

Understanding Why Generalized Reweighting Does Not Improve Over ERM

ICLR 2023poster

Empirical risk minimization (ERM) is known to be non-robust in practice to distributional shift where the training and the test distributions are different. A suite of approaches, such as importance weighting, and variants of distributionally robust optimization (DRO), have been proposed to solve th…

2023

iSCAN: Identifying Causal Mechanism Shifts among Nonlinear Additive Noise Models

NeurIPS 2023poster

Structural causal models (SCMs) are widely used in various disciplines to represent causal relationships among variables in complex systems. Unfortunately, the underlying causal structure is often unknown, and estimating it from data remains a challenging task. In many situations, however, the end…

2022

Analyzing and Improving the Optimization Landscape of Noise-Contrastive Estimation

ICLR 2022spotlight

Noise-contrastive estimation (NCE) is a statistically consistent method for learning unnormalized probabilistic models. It has been empirically observed that the choice of the noise distribution is crucial for NCE’s performance. However, such observation has never been made formal or quantitative. I…

Cited by 22SourcePDFScholar
2022

DAGMA: Learning DAGs via M-matrices and a Log-Determinant Acyclicity Characterization

NeurIPS 2022accept

The combinatorial problem of learning directed acyclic graphs (DAGs) from data was recently framed as a purely continuous optimization problem by leveraging a differentiable acyclicity characterization of DAGs based on the trace of a matrix exponential function. Existing acyclicity characterizations…

2022

FILM: Following Instructions in Language with Modular Methods

ICLR 2022poster

Recent methods for embodied instruction following are typically trained end-to-end using imitation learning. This often requires the use of expert trajectories and low-level language instructions. Such approaches assume that neural states will integrate multimodal semantics to perform state tracking…

2022

First is Better Than Last for Language Data Influence

NeurIPS 2022accept

The ability to identify influential training examples enables us to debug training data and explain model behavior. Existing techniques to do so are based on the flow of training data influence through the model parameters. For large models in NLP applications, it is often computationally infeasible…

2022

Identifiability of deep generative models without auxiliary information

NeurIPS 2022accept

We prove identifiability of a broad class of deep latent variable models that (a) have universal approximation capabilities and (b) are the decoders of variational autoencoders that are commonly used in practice. Unlike existing work, our analysis does not require weak supervision, auxiliary informa…

Cited by 63SourcePDFScholar
2022

Masked Prediction: A Parameter Identifiability View

NeurIPS 2022accept

The vast majority of work in self-supervised learning have focused on assessing recovered features by a chosen set of downstream tasks. While there are several commonly used benchmark datasets, this lens of feature learning requires assumptions on the downstream tasks which are not inherent to the d…

Cited by 9SourcePDFScholar
2021

Boosted CVaR Classification

NeurIPS 2021poster

Many modern machine learning tasks require models with high tail performance, i.e. high performance over the worst-off samples in the dataset. This problem has been widely studied in fields such as algorithmic fairness, class imbalance, and risk-sensitive decision making. A popular approach to maxim…

2021

Evaluations and Methods for Explanation through Robustness Analysis

ICLR 2021poster

Feature based explanations, that provide importance of each feature towards the model prediction, is arguably one of the most intuitive ways to explain a model. In this paper, we establish a novel set of evaluation criteria for such feature based explanations by robustness analysis. In contrast to e…

Cited by 69SourcePDFScholar
2021

Learning latent causal graphs via mixture oracles

NeurIPS 2021poster

We study the problem of reconstructing a causal graphical model from data in the presence of latent variables. The main problem of interest is recovering the causal structure over the latent variables while allowing for general, potentially nonlinear dependencies. In many practical problems, the dep…