← Search

Alexander Wei

7 accepted papers

2024

Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation

ICML 2024poster

Black-box finetuning is an emerging interface for adapting state-of-the-art language models to user needs. However, such access may also let malicious actors undermine model safety. To demonstrate the challenge of defending finetuning interfaces, we introduce covert malicious finetuning, a method to…

Cited by 30SourcePDFScholar
2022

More Than a Toy: Random Matrix Models Predict How Real-World Neural Representations Generalize

ICML 2022spotlight

Of theories for why large-scale machine learning models generalize despite being vastly overparameterized, which of their assumptions are needed to capture the qualitative phenomena of generalization in the real world? On one hand, we find that most theoretical analyses fall short of capturing these…

2022

TCT: Convexifying Federated Learning using Bootstrapped Neural Tangent Kernels

NeurIPS 2022accept

State-of-the-art federated learning methods can perform far worse than their centralized counterparts when clients have dissimilar data distributions. For neural networks, even when centralized SGD easily finds a solution that is simultaneously performant for all clients, current federated optimizat…

2021

Learning Equilibria in Matching Markets from Bandit Feedback

NeurIPS 2021spotlight

Large-scale, two-sided matching platforms must find market outcomes that align with user preferences while simultaneously learning these preferences from data. But since preferences are inherently uncertain during learning, the classical notion of stability (Gale and Shapley, 1962; Shapley and Shubi…

Cited by 49SourcePDFScholar