← Search

Harrison Ngan

1 accepted papers

2025

Representation Bending for Large Language Model Safety

ACL 2025long

Large Language Models (LLMs) have emerged as powerful tools, but their inherent safety risks – ranging from harmful content generation to broader societal harms – pose significant challenges. These risks can be amplified by the recent adversarial attacks, fine-tuning vulnerabilities, and the increas…