← Search

Max Torop

4 accepted papers

2025

DISCO: Disentangled Communication Steering for Large Language Models

NeurIPS 2025poster

A variety of recent methods guide large language model outputs via the inference-time addition of *steering vectors* to residual-stream or attention-head representations. In contrast, we propose to inject steering vectors directly into the query and value representation spaces within attention heads…

Cited by 0SourcecodeScholar
2024

Boundary-Aware Uncertainty for Feature Attribution Explainers

AISTATS 2024poster

Post-hoc explanation methods have become a critical tool for understanding black-box classifiers in high-stakes applications. However, high-performing classifiers are often highly nonlinear and can exhibit complex behavior around the decision boundary, leading to brittle or misleading local explanat…

2023

SmoothHess: ReLU Network Feature Interactions via Stein's Lemma

NeurIPS 2023poster

Several recent methods for interpretability model feature interactions by looking at the Hessian of a neural network. This poses a challenge for ReLU networks, which are piecewise-linear and thus have a zero Hessian almost everywhere. We propose SmoothHess, a method of estimating second-order intera…