← Search

Zirui He

2 accepted papers

2025

SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language Models

EMNLP 2025

Large language models (LLMs) have demonstrated impressive capabilities in natural language understanding and generation, but controlling their behavior reliably remains challenging, especially in open-ended generation settings. This paper introduces a novel supervised steering approach that operates

2024

Mitigating Shortcuts in Language Models with Soft Label Encoding

COLING 2024main

Recent research has shown that large language models rely on spurious correlations in the data for natural language understanding (NLU) tasks. In this work, we aim to answer the following research question: Can we reduce spurious correlations by modifying the ground truth labels of the training data…