← Search

Avani Gupta

4 accepted papers

2025

Building Trust in Clinical LLMs: Bias Analysis and Dataset Transparency

EMNLP 2025

Large language models offer transformative potential for healthcare, yet their responsible and equitable development depends critically on a deeper understanding of how training data characteristics influence model behavior, including the potential for bias. Current practices in dataset curation and

2025

Prototype Guided Backdoor Defense via Activation Space Manipulation

ICCV 2025accepted

Deep learning models are susceptible to backdoor attacks involving malicious perturbation of some training data with a trigger to force misclassification to a target class. Various triggers have been used including semantic triggers that are easily realizable. We present Prototype Guided Backd…

Cited by 0SourcePDFScholar
2023

Concept Distillation: Leveraging Human-Centered Explanations for Model Improvement

NeurIPS 2023poster

Humans use abstract *concepts* for understanding instead of hard features. Recent interpretability research has focused on human-centered concept explanations of neural networks. Concept Activation Vectors (CAVs) estimate a model's sensitivity and possible biases to a given concept. We extend CAVs f…

Cited by 6SourcePDFScholar