← Search

Dong Nguyen

13 accepted papers

2025

Disentangling the Roles of Representation and Selection in Data Pruning

ACL 2025long

Data pruning—selecting small but impactful subsets—offers a promising way to efficiently scale NLP model training. However, existing methods often involve many different design choices, which have not been systematically studied. This limits future developments. In this work, we decompose data pruni…

Cited by 0SourcePDFScholar
2025

FTFT: Efficient and Robust Fine-Tuning by Transferring Training Dynamics

COLING 2025main

Despite the massive success of fine-tuning Pre-trained Language Models (PLMs), they remain susceptible to out-of-distribution input. Dataset cartography is a simple yet effective dual-model approach that improves the robustness of fine-tuned PLMs. It involves fine-tuning a model on the original trai…

2024

What’s Mine becomes Yours: Defining, Annotating and Detecting Context-Dependent Paraphrases in News Interview Dialogs

EMNLP 2024main

Best practices for high conflict conversations like counseling or customer support almost always include recommendations to paraphrase the previous speaker. Although paraphrase classification has received widespread attention in NLP, paraphrases are usually considered independent from context, and c…

2021

Does It Capture STEL? A Modular, Similarity-based Linguistic Style Evaluation Framework

EMNLP 2021main

Style is an integral part of natural language. However, evaluation methods for style measures are rare, often task-specific and usually do not control for content. We propose the modular, fine-grained and content-controlled similarity-based STyle EvaLuation framework (STEL) to test the performance o…

2021

HateCheck: Functional Tests for Hate Speech Detection Models

ACL 2021long

Detecting online hate is a difficult task that even state-of-the-art models struggle with. Typically, hate speech detection models are evaluated by measuring their performance on held-out test data using metrics such as accuracy and F1 score. However, this approach makes it difficult to identify spe…

2021

Introducing CAD: the Contextual Abuse Dataset

NAACL 2021long

Online abuse can inflict harm on users and communities, making online spaces unsafe and toxic. Progress in automatically detecting and classifying abusive content is often held back by the lack of high quality and detailed datasets. We introduce a new dataset of primarily English Reddit entries whic…

2021

On learning and representing social meaning in NLP: a sociolinguistic perspective

NAACL 2021long

The field of NLP has made substantial progress in building meaning representations. However, an important aspect of linguistic meaning, social meaning, has been largely overlooked. We introduce the concept of social meaning to NLP and discuss how insights from sociolinguistics can inform work on rep…

Cited by 42SourcePDFScholar