← Search

Patrick Queiroz Da Silva

1 accepted papers

2025

Steering off Course: Reliability Challenges in Steering Language Models

ACL 2025long

Steering methods for language models (LMs) have gained traction as lightweight alternatives to fine-tuning, enabling targeted modifications to model activations. However, prior studies primarily report results on a few models, leaving critical gaps in understanding the robustness of these methods. I…