← Search

Kush R. Varshney

24 accepted papers

2026

The Shepherd Test: How Will Super Intelligent Agents Balance Care and Control in Asymmetric Relationships?

AAAI 2026technical

This paper introduces the Shepherd Test, a new conceptual test for assessing the moral and relational dimensions of superintelligent artificial agents. The test is inspired by human interactions with animals, where ethical considerations about care, manipulation, and consumption arise in contexts of

Cited by 0SourcePDFScholar
2025

Evaluating the Prompt Steerability of Large Language Models

NAACL 2025long

Building pluralistic AI requires designing models that are able to be shaped to represent a wide range of value systems and cultures. Achieving this requires first being able to evaluate the degree to which a given model is capable of reflecting various personas. To this end, we propose a benchmark…

2025

Granite Guardian: Comprehensive LLM Safeguarding

NAACL 2025industry

The deployment of language models in real-world applications exposes users to various risks, including hallucinations and harmful or unethical content. These challenges highlight the urgent need for robust safeguards to ensure safe and responsible AI. To address this, we introduce Granite Guardian,…

2024

ComVas: Contextual Moral Values Alignment System

IJCAI 2024poster

In contemporary society, the integration of artificial intelligence (AI) systems into various aspects of daily life raises significant ethical concerns. One critical aspect is to ensure that AI systems align with the moral values of the endusers. To that end, we introduce the Contextual Moral Value…

2024

Using Causal Inference to Investigate Contraceptive Discontinuation in Sub-Saharan Africa

IJCAI 2024poster

Discontinuation rates vary by family planning method and across socio-economic contexts. Understanding these variations and their causes is paramount for developing and implementing policies aimed at curbing discontinuation rates. Randomized controlled trials (RCTs) are ideal for obtaining this info…

2024

Value Alignment from Unstructured Text

EMNLP 2024industry

Aligning large language models (LLMs) to value systems has emerged as a significant area of research within the fields of AI and NLP. Currently, this alignment process relies on the availability of high-quality supervised and preference data, which can be both time-consuming and expensive to curate…

Cited by 1SourcePDFScholar
2023

Equi-Tuning: Group Equivariant Fine-Tuning of Pretrained Models

AAAI 2023technical

We introduce equi-tuning, a novel fine-tuning method that transforms (potentially non-equivariant) pretrained models into group equivariant models while incurring minimum L_2 loss between the feature representations of the pretrained and the equivariant models. Large pretrained models can be equi-tu…

Cited by 27SourcePDFScholar
2023

Minimax AUC Fairness: Efficient Algorithm with Provable Convergence

AAAI 2023technical

The use of machine learning models in consequential decision making often exacerbates societal inequity, in particular yielding disparate impact on members of marginalized groups defined by race and gender. The area under the ROC curve (AUC) is widely used to evaluate the performance of a scoring fu…

2023

What Is Missing in IRM Training and Evaluation? Challenges and Solutions

ICLR 2023poster

Invariant risk minimization (IRM) has received increasing attention as a way to acquire environment-agnostic data representations and predictions, and also a principled solution for preventing spurious correlations from being learned and improving models’ out-of-distribution generalization. Yet, rec…

Cited by 8SourcePDFScholar
2022

Fair Infinitesimal Jackknife: Mitigating the Influence of Biased Training Data Points Without Refitting

NeurIPS 2022accept

In consequential decision-making applications, mitigating unwanted biases in machine learning models that yield systematic disadvantage to members of groups delineated by sensitive attributes such as race and gender is one key intervention to strive for equity. Focusing on demographic parity and equ…

Cited by 34SourcePDFScholar
2022

On the Safety of Interpretable Machine Learning: A Maximum Deviation Approach

NeurIPS 2022accept

Interpretable and explainable machine learning has seen a recent surge of interest. We focus on safety as a key motivation behind the surge and make the relationship between interpretability and safety more quantitative. Toward assessing safety, we introduce the concept of *maximum deviation* via an…

Cited by 9SourcePDFScholar
2021

CoFrNets: Interpretable Neural Architecture Inspired by Continued Fractions

NeurIPS 2021poster

In recent years there has been a considerable amount of research on local post hoc explanations for neural networks. However, work on building interpretable neural architectures has been relatively sparse. In this paper, we present a novel neural architecture, CoFrNet, inspired by the form of contin…

Cited by 13SourcePDFScholar
2021

Empirical or Invariant Risk Minimization? A Sample Complexity Perspective

ICLR 2021poster

Recently, invariant risk minimization (IRM) was proposed as a promising solution to address out-of-distribution (OOD) generalization. However, it is unclear when IRM should be preferred over the widely-employed empirical risk minimization (ERM) framework. In this work, we analyze both these framewor…

Cited by 104SourcePDFScholar
2021

Treatment Effect Estimation Using Invariant Risk Minimization

ICASSP 2021accepted

Inferring causal individual treatment effect (ITE) from observational data is a challenging problem whose difficulty is exacerbated by the presence of treatment assignment bias. In this work, we propose a new way to estimate the ITE using the domain generalization framework of invariant risk minimiz…

Cited by 0SourceScholar
2020

Inspection of Blackbox Models for Evaluating Vulnerability in Maternal, Newborn, and Child Health

IJCAI 2020poster

Improving maternal, newborn, and child health (MNCH) outcomes is a critical target for global sustainable development. Our research is centered on building predictive models, evaluating their interpretability, and generating actionable insights about the markers (features) and triggers (events) asso…

Cited by 0SourcePDFScholar
2020

Preservation of Anomalous Subgroups On Variational Autoencoder Transformed Data

ICASSP 2020accepted

We investigate the effect of variational autoencoder (VAE) based data anonymization and its ability to preserve anomalous subgroup properties. We present a Utility Guaranteed Deep Privacy (UGDP) system which casts existing anomalous pattern detection methods as a new utility measure for data synthes…

Cited by 0SourceScholar
2019

Bias Mitigation Post-processing for Individual and Group Fairness

ICASSP 2019accepted

Whereas previous post-processing approaches for increasing the fairness of predictions of biased classifiers address only group fairness, we propose a method for increasing both individual and group fairness. Our novel framework includes an individual bias detector used to prioritize data samples in…

Cited by 0SourceScholar
2019

Constructing and Compressing Frames in Blockchain-based Verifiable Multi-party Computation

ICASSP 2019accepted

In previous work, we proposed a scalable multi-party verification scheme for expensive iterative computations on a Blockchain substrate by appropriate storage and endorsement of frames of iterates. In this work, we extend the framework to verify sets of complete computations with different unordered…

Cited by 0SourceScholar
2017

Optimized Pre-Processing for Discrimination Prevention

NeurIPS 2017poster

Non-discrimination is a recognized objective in algorithmic decision making. In this paper, we introduce a novel probabilistic formulation of data pre-processing for reducing discrimination. We propose a convex optimization for learning a data transformation with three goals: controlling discriminat…

Cited by 1131SourcePDFScholar
2015

Learning interpretable classification rules using sequential rowsampling

ICASSP 2015accepted

In our previous work we have presented an approach to learn interpretable classification rules using a Boolean compressed sensing formulation. Our approach uses a linear programming (LP) relaxation and allows us to find interpretable (sparse) classification rules that achieve good generalization acc…

Cited by 0SourceScholar