← Search

Nika Haghtalab

29 accepted papers

2026

Distortion of AI Alignment Revisited: RLHF is a Decent Utilitarian Aligner

ICML 2026poster

While Reinforcement Learning from Human Feedback (RLHF) is the standard paradigm for aligning large language models with human preferences, its effectiveness in pluralistic settings has been called into question. Notably, recent work by Golz et al. (2025) demonstrated that the *distortion* — defined…

Cited by 0SourceScholar
2026

Subliminal Effects in Your Data: A General Mechanism via Log-Linearity

ICML 2026poster

Training modern large language models (LLMs) has become a veritable smorgasbord of algorithms and datasets designed to elicit particular behaviors, making it critical to develop techniques to understand the effects of datasets on the model's properties. This is exacerbated by recent experiments that…

Cited by 0SourceScholar
2026

Three Years of r/ChatGPT: Societal Impact Evaluations from Social Media Data

ICML 2026poster

ChatGPT was launched on November 30, 2022; the r/ChatGPT subreddit was created just one day later. Since then, chatbot-based AI products have gone from niche proofs-of-concept to widely-used household names. However, the ways in which adoption has developed, especially among non-experts, remains poo…

Cited by 0SourceScholar
2025

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?

NeurIPS 2025poster

After pre-training, large language models are aligned with human preferences based on pairwise comparisons. State-of-the-art alignment methods (such as PPO-based RLHF and DPO) are built on the assumption of aligning with a single preference model, despite being deployed in settings where users have…

Cited by 0SourceScholar
2025

From Style to Facts: Mapping the Boundaries of Knowledge Injection with Finetuning

NeurIPS 2025poster

Finetuning provides a scalable and cost-effective means of customizing language models for specific tasks or response styles, with greater reliability than prompting or in-context learning. In contrast, the conventional wisdom is that injecting knowledge via finetuning results in brittle performance…

Cited by 0SourceScholar
2024

Can Probabilistic Feedback Drive User Impacts in Online Platforms?

AISTATS 2024poster

A common explanation for negative user impacts of content recommender systems is misalignment between the platform’s objective and user welfare. In this work, we show that misalignment in the platform’s objective is not the only potential cause of unintended impacts on users: even when the platform’…

Cited by 8SourcePDFScholar
2024

Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation

ICML 2024poster

Black-box finetuning is an emerging interface for adapting state-of-the-art language models to user needs. However, such access may also let malicious actors undermine model safety. To demonstrate the challenge of defending finetuning interfaces, we introduce covert malicious finetuning, a method to…

Cited by 30SourcePDFScholar
2024

Delegating Data Collection in Decentralized Machine Learning

AISTATS 2024poster

Motivated by the emergence of decentralized machine learning (ML) ecosystems, we study the delegation of data collection. Taking the field of contract theory as our starting point, we design optimal and near-optimal contracts that deal with two fundamental information asymmetries that arise in decen…

Cited by 11SourcePDFScholar
2024

Is Knowledge Power? On the (Im)possibility of Learning from Strategic Interactions

NeurIPS 2024poster

When learning in strategic environments, a key question is whether agents can overcome uncertainty about their preferences to achieve outcomes they could have achieved absent any uncertainty. Can they do this solely through interactions with each other? We focus this question on the ability of agent…

Cited by 4SourcePDFScholar
2023

A Unifying Perspective on Multi-Calibration: Game Dynamics for Multi-Objective Learning

NeurIPS 2023poster

We provide a unifying framework for the design and analysis of multi-calibrated predictors. By placing the multi-calibration problem in the general setting of multi-objective learning---where learning guarantees must hold simultaneously over a set of distributions and loss functions---we exploit con…

Cited by 16SourcePDFScholar
2023

Calibrated Stackelberg Games: Learning Optimal Commitments Against Calibrated Agents

NeurIPS 2023spotlight

In this paper, we introduce a generalization of the standard Stackelberg Games (SGs) framework: _Calibrated Stackelberg Games_. In CSGs, a principal repeatedly interacts with an agent who (contrary to standard SGs) does not have direct access to the principal's action but instead best responds to _c…

Cited by 34SourcePDFScholar
2023

Competition, Alignment, and Equilibria in Digital Marketplaces

AAAI 2023technical

Competition between traditional platforms is known to improve user utility by aligning the platform's actions with user preferences. But to what extent is alignment exhibited in data-driven marketplaces? To study this question from a theoretical perspective, we introduce a duopoly market where platf…

Cited by 20SourcePDFScholar
2023

Improved Bayes Risk Can Yield Reduced Social Welfare Under Competition

NeurIPS 2023poster

As the scale of machine learning models increases, trends such as scaling laws anticipate consistent downstream improvements in predictive accuracy. However, these trends take the perspective of a single model-provider in isolation, while in reality providers often compete with each other for users.…

2022

On-Demand Sampling: Learning Optimally from Multiple Distributions

NeurIPS 2022accept

Societal and real-world considerations such as robustness, fairness, social welfare and multi-agent tradeoffs have given rise to multi-distribution learning paradigms, such as collaborative [Blum et al. 2017], group distributionally robust [Sagawa et al. 2019], and fair federated learning [Mohri et…

2022

Oracle-Efficient Online Learning for Smoothed Adversaries

NeurIPS 2022accept

We study the design of computationally efficient online learning algorithms under smoothed analysis. In this setting, at every step, an adversary generates a sample from an adaptively chosen distribution whose density is upper bounded by $1/\sigma$ times the uniform density. Given access to an offli…

Cited by 14SourcePDFScholar
2021

One for One, or All for All: Equilibria and Optimality of Collaboration in Federated Learning

ICML 2021spotlight

In recent years, federated learning has been embraced as an approach for bringing about collaboration across large populations of learning agents. However, little is known about how collaboration protocols should take agents’ incentives into account when allocating individual resources for communal…

2020

Maximizing Welfare with Incentive-Aware Evaluation Mechanisms

IJCAI 2020poster

Motivated by applications such as college admission and insurance rate determination, we study a classification problem where the inputs are controlled by strategic individuals who can modify their features at a cost. A learner can only partially observe the features, and aims to classify individua…

Cited by 0SourcePDFScholar
2020

Smoothed Analysis of Online and Differentially Private Learning

NeurIPS 2020spotlight

Practical and pervasive needs for robustness and privacy in algorithms have inspired the design of online adversarial and differentially private learning algorithms. The primary quantity that characterizes learnability in these settings is the Littlestone dimension of the class of hypotheses [Ben-Da…

Cited by 69SourcePDFScholar
2019

Structured Robust Submodular Maximization: Offline and Online Algorithms

AISTATS 2019poster

Constrained submodular function maximization has been used in subset selection problems such as selection of most informative sensor locations. While these models have been quite popular, the solutions obtained via this approach are unstable to perturbations in data defining the submodular functions…

Cited by 42SourcePDFScholar
2019

Toward a Characterization of Loss Functions for Distribution Learning

NeurIPS 2019poster

In this work we study loss functions for learning and evaluating probability distributions over large discrete domains. Unlike classification or regression where a wide variety of loss functions are used, in the distribution learning and density estimation literature, very few losses outside the do…

Cited by 8SourcePDFScholar