← Search

Nils Feldhus

6 accepted papers

2025

Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description Framework

NeurIPS 2025poster

Automated interpretability research aims to identify concepts encoded in neural network features to enhance human understanding of model behavior. Within the context of large language models (LLMs) for natural language processing (NLP), current automated neuron-level feature description methods face…

Cited by 0SourceScholar
2025

Cross-Refine: Improving Natural Language Explanation Generation by Learning in Tandem

COLING 2025main

Natural language explanations (NLEs) are vital for elucidating the reasoning behind large language model (LLM) decisions. Many techniques have been developed to generate NLEs using LLMs. However, like humans, LLMs might not always produce optimal NLEs on first attempt. Inspired by human learning pro…

2025

FitCF: A Framework for Automatic Feature Importance-guided Counterfactual Example Generation

ACL 2025finding

Counterfactual examples are widely used in natural language processing (NLP) as valuable data to improve models, and in explainable artificial intelligence (XAI) to understand model behavior. The automated generation of counterfactual examples remains a challenging task even for large language model…

2024

CoXQL: A Dataset for Parsing Explanation Requests in Conversational XAI Systems

EMNLP 2024finding

Conversational explainable artificial intelligence (ConvXAI) systems based on large language models (LLMs) have garnered significant interest from the research community in natural language processing (NLP) and human-computer interaction (HCI). Such systems can provide answers to user questions abou…

2023

InterroLang: Exploring NLP Models and Datasets through Dialogue-based Explanations

EMNLP 2023long findings

While recently developed NLP explainability methods let us open the black box in various ways (Madsen et al., 2022), a missing ingredient in this endeavor is an interactive tool offering a conversational interface. Such a dialogue system can help users explore datasets and models with explanations i…

Cited by 0SourcecodeScholar
2021

Thermostat: A Large Collection of NLP Model Explanations and Analysis Tools

EMNLP 2021system demonstrations

In the language domain, as in other domains, neural explainability takes an ever more important role, with feature attribution methods on the forefront. Many such methods require considerable computational resources and expert knowledge about implementation details and parameter choices. To facilita…