← Search

Trenton Chang

4 accepted papers

2026

A Course Correction in Steerability Evaluation: Revealing Miscalibration and Side Effects in LLMs

AAAI 2026technical

Despite advances in large language models (LLMs) on reasoning and instruction-following benchmarks, it is unclear whether they can reliably produce outputs aligned with a variety of user goals, a concept called steerability. We highlight two gaps in current LLM evaluations for assessing steerability

Cited by 0SourcePDFScholar
2025

Disentangling misreporting from genuine adaptation in strategic settings: a causal approach

NeurIPS 2025poster

In settings where ML models are used to inform the allocation of resources, agents affected by the allocation decisions might have an incentive to strategically change their features to secure better outcomes. While prior work has studied strategic responses broadly, disentangling misreporting from…

Cited by 0SourceScholar
2024

From Biased Selective Labels to Pseudo-Labels: An Expectation-Maximization Framework for Learning from Biased Decisions

ICML 2024poster

Selective labels occur when label observations are subject to a decision-making process; e.g., diagnoses that depend on the administration of laboratory tests. We study a clinically-inspired selective label problem called disparate censorship, where labeling biases vary across subgroups and unlabele…

2024

Who’s Gaming the System? A Causally-Motivated Approach for Detecting Strategic Adaptation

NeurIPS 2024poster

In many settings, machine learning models may be used to inform decisions that impact individuals or entities who interact with the model. Such entities, or *agents,* may *game* model decisions by manipulating their inputs to the model to obtain better outcomes and maximize some utility. We consider…