← Search

Marika Swanberg

4 accepted papers

2025

Measuring memorization in language models via probabilistic extraction

NAACL 2025long

Large language models (LLMs) are susceptible to memorizing training data, raising concerns about the potential extraction of sensitive information at generation time. Discoverable extraction is the most common method for measuring this issue: split a training example into a prefix and suffix, then p…

Cited by 2SourcePDFScholar
2025

Privacy in Metalearning and Multitask Learning: Modeling and Separations

AISTATS 2025poster

Model personalization allows a set of individuals, each facing a different learning task, to train models that are more accurate for each person than those they could develop individually. The goals of personalization are captured in a variety of formal frameworks, such as multitask learning and met…

Cited by 0SourceScholar
2024

Auditing Privacy Mechanisms via Label Inference Attacks

NeurIPS 2024spotlight

We propose reconstruction advantage measures to audit label privatization mechanisms. A reconstruction advantage measure quantifies the increase in an attacker's ability to infer the true label of an unlabeled example when provided with a private version of the labels in a dataset (e.g., aggregate o…

Cited by 1SourcePDFScholar
2021

Differentially Private Sampling from Distributions

NeurIPS 2021poster

We initiate an investigation of private sampling from distributions. Given a dataset with $n$ independent observations from an unknown distribution $P$, a sampling algorithm must output a single observation from a distribution that is close in total variation distance to $P$ while satisfying differ…

Cited by 12SourcePDFScholar