← Search

Jason Wang

3 accepted papers

2023

MoPe: Model Perturbation based Privacy Attacks on Language Models

EMNLP 2023long main

Recent work has shown that Large Language Models (LLMs) can unintentionally leak sensitive information present in their training data. In this paper, we present Model Perturbations (MoPe), a new method to identify with high confidence if a given text is in the training data of a pre-trained languag…

Cited by 33SourceScholar
2023

The Naughtyformer: A Transformer Understands and Moderates Adult Humor (Student Abstract)

AAAI 2023technical

Jokes are intentionally written to be funny, but not all jokes are created the same. While recent work has shown impressive results on humor detection in text, we instead investigate the more nuanced task of detecting humor subtypes, especially of the more adult variety. To that end, we introduce a…

Cited by 1SourcePDFScholar