← Search

Asma Ghandeharioun

9 accepted papers

2025

Racing Thoughts: Explaining Contextualization Errors in Large Language Models

NAACL 2025long

The profound success of transformer-based language models can largely be attributed to their ability to integrate relevant contextual information from an input sequence in order to generate a response or complete a task. However, we know very little about the algorithms that a model employs to imple…

Cited by 0SourcePDFScholar
2024

Interpretability Illusions in the Generalization of Simplified Models

ICML 2024poster

A common method to study deep learning systems is to use simplified model representations—for example, using singular value decomposition to visualize the model’s hidden states in a lower dimensional space. This approach assumes that the results of these simplifications are faithful to the original…

Cited by 14SourcePDFScholar
2024

Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models

ICML 2024poster

Understanding the internal representations of large language models (LLMs) can help explain models' behavior and verify their alignment with human values. Given the capabilities of LLMs in generating human-understandable text, we propose leveraging the model itself to explain its internal representa…

Cited by 64SourcePDFScholar
2024

Who's asking? User personas and the mechanics of latent misalignment

NeurIPS 2024spotlight

Studies show that safety-tuned models may nevertheless divulge harmful information. In this work, we show that whether they do so depends significantly on who they are talking to, which we refer to as *user persona*. In fact, we find manipulating user persona to be more effective for eliciting harmf…

Cited by 5SourcePDFScholar
2023

Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language Models

NeurIPS 2023spotlight

Language models learn a great quantity of factual information during pretraining, and recent work localizes this information to specific model weights like mid-layer MLP weights. In this paper, we find that we can change how a fact is stored in a model by editing weights that are in a different loca…

2023

Post Hoc Explanations of Language Models Can Improve Language Models

NeurIPS 2023poster

Large Language Models (LLMs) have demonstrated remarkable capabilities in performing complex tasks. Moreover, recent research has shown that incorporating human-annotated rationales (e.g., Chain-of-Thought prompting) during in-context learning can significantly enhance the performance of these model…

Cited by 72SourcePDFScholar
2022

DISSECT: Disentangled Simultaneous Explanations via Concept Traversals

ICLR 2022poster

Explaining deep learning model inferences is a promising venue for scientific understanding, improving safety, uncovering hidden biases, evaluating fairness, and beyond, as argued by many scholars. One of the principal benefits of counterfactual explanations is allowing users to explore "what-if" sc…

2019

Approximating Interactive Human Evaluation with Self-Play for Open-Domain Dialog Systems

NeurIPS 2019poster

Building an open-domain conversational agent is a challenging problem. Current evaluation methods, mostly post-hoc judgments of static conversation, do not capture conversation quality in a realistic interactive context. In this paper, we investigate interactive human evaluation and provide evidence…

2018

Multimodal Prediction and Personalization of Photo Edits with Deep Generative Models

AISTATS 2018poster

Professional-grade software applications are powerful but complicated – expert users can achieve impressive results, but novices often struggle to complete even basic tasks. Photo editing is a prime example: after loading a photo, the user is confronted with an array of cryptic sliders like "clarity…

Cited by 0SourcePDFScholar