← Search

Aaron Jiaxun Li

2 accepted papers

2025

More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness

ICLR 2025oral

The trustworthiness of Large Language Models (LLMs) refers to the extent to which their outputs are reliable, safe, and ethically aligned, and it has become a crucial consideration alongside their cognitive performance. In practice, Reinforcement Learning From Human Feedback (RLHF) has been widely u…

2024

Improving Prototypical Visual Explanations with Reward Reweighing, Reselection, and Retraining

ICML 2024poster

In recent years, work has gone into developing deep interpretable methods for image classification that clearly attributes a model's output to specific features of the data. One such of these methods is the Prototypical Part Network (ProtoPNet), which attempts to classify images based on meaningful…