← Search

Eoin Delaney

4 accepted papers

2026

The Heterogeneous Safety Impacts of Benign Multilingual Fine-Tuning

ICML 2026poster

Fine-tuning a large language model is a ubiquitous method for enhancing its capability on a specific downstream task. However, prior work has shown that this increase in capability comes with a cost: it can increase a model's tendency to respond to unsafe adversarial prompts, even when fine-tuning w…

Cited by 0SourceScholar
2024

Counterfactual Explanations for Misclassified Images: How Human and Machine Explanations Differ (Abstract Reprint)

AAAI 2024technical

Counterfactual explanations have emerged as a popular solution for the eXplainable AI (XAI) problem of elucidating the predictions of black-box deep-learning systems because people easily understand them, they apply across different problem domains and seem to be legally compliant. Although over 100…

Cited by 0SourcePDFScholar
2023

Advancing Post-Hoc Case-Based Explanation with Feature Highlighting

IJCAI 2023poster

Explainable AI (XAI) has been proposed as a valuable tool to assist in downstream tasks involving human-AI collaboration. Perhaps the most psychologically valid XAI techniques are case-based approaches which display "whole" exemplars to explain the predictions of black-box AI systems. However, for s…

2021

If Only We Had Better Counterfactual Explanations: Five Key Deficits to Rectify in the Evaluation of Counterfactual XAI Techniques

IJCAI 2021poster

In recent years, there has been an explosion of AI research on counterfactual explanations as a solution to the problem of eXplainable AI (XAI). These explanations seem to offer technical, psychological and legal benefits over other explanation techniques. We survey 100 distinct counterfactual expl…

Cited by 206SourcePDFScholar