← Search

Frank Rudzicz

29 accepted papers

2025

ACCORD: Closing the Commonsense Measurability Gap

NAACL 2025long

We present ACCORD, a framework and benchmark suite for disentangling the commonsense grounding and reasoning abilities of large language models (LLMs) through controlled, multi-hop counterfactuals. ACCORD introduces formal elements to commonsense reasoning to explicitly control and quantify reasonin…

Cited by 0SourcePDFScholar
2025

Filtered not Mixed: Filtering-Based Online Gating for Mixture of Large Language Models

ICLR 2025poster

We propose MoE-F — a formalized mechanism for combining N pre-trained expert Large Language Models (LLMs) in online time-series prediction tasks by adaptively forecasting the best weighting of LLM predictions at every time step. Our mechanism leverages the conditional information in each expert's ru…

2025

Not Lost After All: How Cross-Encoder Attribution Challenges Position Bias Assumptions in LLM Summarization

EMNLP 2025

Position bias, the tendency of Large Language Models (LLMs) to select content based on its structural position in a document rather than its semantic relevance, has been viewed as a key limitation in automatic summarization. To measure position bias, prior studies rely heavily on n-gram matching tec

Cited by 0SourcePDFScholar
2025

Trustworthy Medical Question Answering: An Evaluation-Centric Survey

EMNLP 2025

Trustworthiness in healthcare question-answering (QA) systems is important for ensuring patient safety, clinical effectiveness, and user confidence. As large language models (LLMs) become increasingly integrated into medical settings, the reliability of their responses directly influences clinical d

Cited by 0SourcePDFScholar
2024

Auxiliary Knowledge-Induced Learning for Automatic Multi-Label Medical Document Classification

COLING 2024main

The International Classification of Diseases (ICD) is an authoritative medical classification system of different diseases and conditions for clinical and management purposes. ICD indexing aims to assign a subset of ICD codes to a medical record. Since human coding is labour-intensive and error-pron…

Cited by 0SourcePDFScholar
2024

Graph-tree Fusion Model with Bidirectional Information Propagation for Long Document Classification

EMNLP 2024finding

Long document classification presents challenges in capturing both local and global dependencies due to their extensive content and complex structure. Existing methods often struggle with token limits and fail to adequately model hierarchical relationships within documents. To address these constrai…

Cited by 0SourcePDFScholar
2024

Immunization against harmful fine-tuning attacks

EMNLP 2024finding

Large Language Models (LLMs) are often trained with safety guards intended to prevent harmful text generation. However, such safety training can be removed by fine-tuning the LLM on harmful datasets. While this emerging threat (harmful fine-tuning attacks) has been characterized by previous work, th…

Cited by 20SourcePDFScholar
2024

Long-form evaluation of model editing

NAACL 2024long

Evaluations of model editing, a technique for changing the factual knowledge held by Large Language Models (LLMs), currently only use the ‘next few token’ completions after a prompt. As a result, the impact of these methods on longer natural language generation is largely unknown. We introduce long-…

2024

Multi-stage Retrieve and Re-rank Model for Automatic Medical Coding Recommendation

NAACL 2024long

The International Classification of Diseases (ICD) serves as a definitive medical classification system encompassing a wide range of diseases and conditions. The primary objective of ICD indexing is to allocate a subset of ICD codes to a medical record, which facilitates standardized documentation a…

Cited by 2SourcePDFScholar
2024

Representation Noising: A Defence Mechanism Against Harmful Finetuning

NeurIPS 2024poster

Releasing open-source large language models (LLMs) presents a dual-use risk since bad actors can easily fine-tune these models for harmful purposes. Even without the open release of weights, weight stealing and fine-tuning APIs make closed models vulnerable to harmful fine-tuning attacks (HFAs). Whi…

2023

Improving Automatic Quotation Attribution in Literary Novels

ACL 2023short

Current models for quotation attribution in literary novels assume varying levels of available information in their training and test data, which poses a challenge for in-the-wild inference. Here, we approach quotation attribution as a set of four interconnected sub-tasks: character identification,…

2022

A Remedy For Distributional Shifts Through Expected Domain Translation

ICASSP 2022accepted

Machine learning models often fail to generalize to unseen domains due to the distributional shifts. A family of such shifts, “correlation shifts,” is caused by spurious correlations in the data. It is studied under the overarching topic of “domain generalization.” In thi…

Cited by 0SourceScholar
2022

Neural reality of argument structure constructions

ACL 2022long

In lexicalist linguistic theories, argument structure is assumed to be predictable from the meaning of verbs. As a result, the verb is the primary determinant of the meaning of a clause. In contrast, construction grammarians propose that argument structure is encoded in constructions (or form-meanin…

2021

An unsupervised framework for tracing textual sources of moral change

EMNLP 2021finding

Morality plays an important role in social well-being, but people’s moral perception is not stable and changes over time. Recent advances in natural language processing have shown that text is an effective medium for informing moral change, but no attempt has been made to quantify the origins of the…

2021

Coughwatch: Real-World Cough Detection using Smartwatches

ICASSP 2021accepted

Continuous monitoring of cough may provide insights into the health of individuals as well as the effectiveness of treatments. Smart-watches, in particular, are highly promising for such monitoring: they are inexpensive, unobtrusive, programmable, and have a variety of sensors. However, current mobi…

Cited by 0SourceScholar
2021

Grad2Task: Improved Few-shot Text Classification Using Gradients for Task Representation

NeurIPS 2021poster

Large pretrained language models (LMs) like BERT have improved performance in many disparate natural language processing (NLP) tasks. However, fine tuning such models requires a large number of training examples for each target task. Simultaneously, many realistic NLP problems are "few shot", withou…

2021

How is BERT surprised? Layerwise detection of linguistic anomalies

ACL 2021long

Transformer language models have shown remarkable ability in detecting when a word is anomalous in context, but likelihood scores offer no information about the cause of the anomaly. In this work, we use Gaussian models for density estimation at intermediate layers of three language models (BERT, Ro…

2020

Speaker Diarization with Session-Level Speaker Embedding Refinement Using Graph Neural Networks

ICASSP 2020accepted

Deep speaker embedding models have been commonly used as a building block for speaker diarization systems; however, the speaker embedding model is usually trained according to a global loss defined on the training data, which could be suboptimal for distinguishing speakers locally in a specific meet…

Cited by 0SourceScholar
2019

Centroid-based Deep Metric Learning for Speaker Recognition

ICASSP 2019accepted

Speaker embedding models that utilize neural networks to map utterances to a space where distances reflect similarity between speakers have driven recent progress in the speaker recognition task. However, there is still a significant performance gap between recognizing speakers in the training set a…

Cited by 0SourceScholar
2017

Multi-view representation learning via gcca for multimodal analysis of Parkinson's disease

ICASSP 2017accepted

Information from different bio-signals such as speech, handwriting, and gait have been used to monitor the state of Parkinson's disease (PD) patients, however, all the multimodal bio-signals may not always be available. We propose a method based on multi-view representation learning via generalized…

Cited by 35SourceScholar
2017

On the impact of non-modal phonation on phonological features

ICASSP 2017accepted

Different modes of vibration of the vocal folds contribute significantly to the voice quality. The neutral mode phonation, often used in a modal voice, is one against which the other modes can be contrastively described, also called non-modal phonations. This paper investigates the impact of non-mod…

Cited by 0SourceScholar
2015

EEG dimensionality reduction in automatic identification of synonymy

ICASSP 2015accepted

Recent work has demonstrated the feasibility of extracting semantic categories directly from cortical measures (e.g., electroencephalography, EEG) during receptive tasks. Here, we automatically classify speech stimuli as either synonymous or non-synonymous with a prior prime in a speech-receptive ta…

Cited by 0SourceScholar