← Search

Vishal Pramanik

3 accepted papers

2026

Hessian-Enhanced Token Attribution (HETA): Interpreting Autoregressive LLMs

ICLR 2026poster

Attribution methods seek to explain language model predictions by quantifying the contribution of input tokens to generated outputs. However, most existing techniques are designed for encoder-based architectures and rely on linear approximations that fail to capture the causal and semantic complexit…

Cited by 1SourcecodeScholar
2026

Jailbreaking the Matrix: Nullspace Steering for Controlled Model Subversion

ICLR 2026poster

Large language models remain vulnerable to jailbreak attacks, inputs crafted to bypass safety mechanisms and elicit harmful responses, despite advances in alignment and instruction tuning. Existing attacks often rely on prompt rewrites, dense optimization, or ad hoc heuristics, and lack interpretabi…

Cited by 0SourcecodeScholar
2026

Selective Amnesia using Contrastive Subnet Erasure for Class Level Unlearning in Vision Models

CVPR 2026

We study concept-level forgetting in pretrained vision models: removing an entire semantic category so the system no longer recognizes that object in unseen images and contexts, rather than merely forgetting specific training examples. Prior work either applies blunt global projections or fine-tunes

Cited by 0SourcecodeScholar