← Search

Hassan Sajjad

25 accepted papers

2025

Data-centric Prediction Explanation via Kernelized Stein Discrepancy

ICLR 2025poster

Existing example-based prediction explanation methods often bridge test and training data points through the model’s parameters or latent representations. While these methods offer clues to the causes of model predictions, they often exhibit innate shortcomings, such as incurring significant computa…

2025

Dependency Parsing is More Parameter-Efficient with Normalization

NeurIPS 2025poster

Dependency parsing is the task of inferring natural language structure, often approached by modeling word interactions via attention through biaffine scoring. This mechanism works like self-attention in Transformers, where scores are calculated for every pair of words in a sentence. However, unlike…

Cited by 0SourcecodeScholar
2025

Explaining the role of Intrinsic Dimensionality in Adversarial Training

ICML 2025poster

Adversarial Training (AT) impacts different architectures in distinct ways: vision models gain robustness but face reduced generalization, encoder-based models exhibit limited robustness improvements with minimal generalization loss, and recent work in latent-space adversarial training demonstrates…

Cited by 0SourcePDFScholar
2025

Not Lost After All: How Cross-Encoder Attribution Challenges Position Bias Assumptions in LLM Summarization

EMNLP 2025

Position bias, the tendency of Large Language Models (LLMs) to select content based on its structural position in a document rather than its semantic relevance, has been viewed as a key limitation in automatic summarization. To measure position bias, prior studies rely heavily on n-gram matching tec

Cited by 0SourcePDFScholar
2024

Immunization against harmful fine-tuning attacks

EMNLP 2024finding

Large Language Models (LLMs) are often trained with safety guards intended to prevent harmful text generation. However, such safety training can be removed by fine-tuning the LLM on harmful datasets. While this emerging threat (harmful fine-tuning attacks) has been characterized by previous work, th…

Cited by 20SourcePDFScholar
2024

Latent Concept-based Explanation of NLP Models

EMNLP 2024main

Interpreting and understanding the predictions made by deep learning models poses a formidable challenge due to their inherently opaque nature. Many previous efforts aimed at explaining these predictions rely on input features, specifically, the words within NLP models. However, such explanations ar…

2024

Long-form evaluation of model editing

NAACL 2024long

Evaluations of model editing, a technique for changing the factual knowledge held by Large Language Models (LLMs), currently only use the ‘next few token’ completions after a prompt. As a result, the impact of these methods on longer natural language generation is largely unknown. We introduce long-…

2024

Multilingual Nonce Dependency Treebanks: Understanding how Language Models Represent and Process Syntactic Structure

NAACL 2024long

We introduce SPUD (Semantically Perturbed Universal Dependencies), a framework for creating nonce treebanks for the multilingual Universal Dependencies (UD) corpora. SPUD data satisfies syntactic argument structure, provides syntactic annotations, and ensures grammaticality via language-specific rul…

2024

Representation Noising: A Defence Mechanism Against Harmful Finetuning

NeurIPS 2024poster

Releasing open-source large language models (LLMs) presents a dual-use risk since bad actors can easily fine-tune these models for harmful purposes. Even without the open release of weights, weight stealing and fine-tuning APIs make closed models vulnerable to harmful fine-tuning attacks (HFAs). Whi…

2024

SUGARCREPE++ Dataset: Vision-Language Model Sensitivity to Semantic and Lexical Alterations

NeurIPS 2024poster

Despite their remarkable successes, state-of-the-art large language models (LLMs), including vision-and-language models (VLMs) and unimodal language models (ULMs), fail to understand precise semantics. For example, semantically equivalent sentences expressed using different lexical compositions elic…

2023

ConceptX: A Framework for Latent Concept Analysis

AAAI 2023technical

The opacity of deep neural networks remains a challenge in deploying solutions where explanation is as important as precision. We present ConceptX, a human-in-the-loop framework for interpreting and annotating latent representational space in pre-trained Language Models (pLMs). We use an unsupervise…

2023

Impact of Adversarial Training on Robustness and Generalizability of Language Models

ACL 2023findings

Adversarial training is widely acknowledged as the most effective defense against adversarial attacks. However, it is also well established that achieving both robustness and generalization in adversarially trained models involves a trade-off. The goal of this work is to provide an in depth comparis…

Cited by 9SourcePDFScholar
2022

Analyzing Encoded Concepts in Transformer Language Models

NAACL 2022long

We propose a novel framework ConceptX, to analyze how latent concepts are encoded in representations learned within pre-trained lan-guage models. It uses clustering to discover the encoded concepts and explains them by aligning with a large set of human-defined concepts. Our analysis on seven transf…

2022

Discovering Latent Concepts Learned in BERT

ICLR 2022poster

A large number of studies that analyze deep neural network models and their ability to encode various linguistic and non-linguistic concepts provide an interpretation of the inner mechanics of these models. The scope of the analyses is limited to pre-defined concepts that reinforce the traditional l…

Cited by 77SourcePDFScholar
2022

Effect of Post-processing on Contextualized Word Representations

COLING 2022main

Post-processing of static embedding has been shown to improve their performance on both lexical and sequence-level tasks. However, post-processing for contextualized embeddings is an under-studied problem. In this work, we question the usefulness of post-processing for contextualized embeddings obta…

Cited by 17SourcePDFScholar
2022

On the Transformation of Latent Space in Fine-Tuned NLP Models

EMNLP 2022main

We study the evolution of latent space in fine-tuned NLP models. Different from the commonly used probing-framework, we opt for an unsupervised method to analyze representations. More specifically, we discover latent concepts in the representational space using hierarchical clustering. We then use a…

Cited by 23SourcePDFScholar
2022

Probing for Constituency Structure in Neural Language Models

EMNLP 2022finding

In this paper, we investigate to which extent contextual neural language models (LMs) implicitly learn syntactic structure. More concretely, we focus on constituent structure as represented in the Penn Treebank (PTB). Using standard probing techniques based on diagnostic classifiers, we assess the a…

2021

Fighting the COVID-19 Infodemic: Modeling the Perspective of Journalists, Fact-Checkers, Social Media Platforms, Policy Makers, and the Society

EMNLP 2021finding

With the emergence of the COVID-19 pandemic, the political and the medical aspects of disinformation merged as the problem got elevated to a whole new level to become the first global infodemic. Fighting this infodemic has been declared one of the most important focus areas of the World Health Organ…

2020

AraBench: Benchmarking Dialectal Arabic-English Machine Translation

COLING 2020main

Low-resource machine translation suffers from the scarcity of training data and the unavailability of standard evaluation sets. While a number of research efforts target the former, the unavailability of evaluation benchmarks remain a major hindrance in tracking the progress in low-resource machine…

Cited by 39SourcePDFScholar
2020

Are We Ready for this Disaster? Towards Location Mention Recognition from Crisis Tweets

COLING 2020main

The widespread usage of Twitter during emergencies has provided a new opportunity and timely resource to crisis responders for various disaster management tasks. Geolocation information of pertinent tweets is crucial for gaining situational awareness and delivering aid. However, the majority of twee…

Cited by 16SourcePDFScholar
2019

Identifying and Controlling Important Neurons in Neural Machine Translation

ICLR 2019poster

Neural machine translation (NMT) models learn representations containing substantial linguistic information. However, it is not clear if such information is fully distributed or if some of it can be attributed to individual neurons. We develop unsupervised methods for discovering important neurons i…

Cited by 218SourcePDFScholar