← Search

Natalie Mackraz

5 accepted papers

2025

Aligning LLMs by Predicting Preferences from User Writing Samples

ICML 2025poster

Accommodating human preferences is essential for creating aligned LLM agents that deliver personalized and effective interactions. Recent work has shown the potential for LLMs acting as writing agents to infer a description of user preferences. Agent alignment then comes from conditioning on the inf…

2025

Bias after Prompting: Persistent Discrimination in Large Language Models

EMNLP 2025

A dangerous assumption that can be made from prior work on the bias transfer hypothesis (BTH) is that biases do not transfer from pre-trained large language models (LLMs) to adapted models. We invalidate this assumption by studying the BTH in causal models under prompt adaptations, as prompting is a

Cited by 0SourcePDFScholar
2025

Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs

ICML 2025poster

The recent rapid adoption of large language models (LLMs) highlights the critical need for benchmarking their fairness. Conventional fairness metrics, which focus on discrete accuracy-based evaluations (i.e., prediction correctness), fail to capture the implicit impact of model uncertainty (e.g., hi…

2024

Large Language Models as Generalizable Policies for Embodied Tasks

ICLR 2024poster

We show that large language models (LLMs) can be adapted to be generalizable policies for embodied visual tasks. Our approach, called Large LAnguage model Reinforcement Learning Policy (LLaRP), adapts a pre-trained frozen LLM to take as input text instructions and visual egocentric observations and…

Cited by 75SourcePDFScholar
2023

Sample-Efficient Preference-based Reinforcement Learning with Dynamics Aware Rewards

CoRL 2023poster

Preference-based reinforcement learning (PbRL) aligns a robot behavior with human preferences via a reward function learned from binary feedback over agent behaviors. We show that encoding environment dynamics in the reward function improves the sample efficiency of PbRL by an order of magnitude. In…

Cited by 7SourcecodeScholar