← Search

Sumit Kumar Jha

7 accepted papers

2026

Hessian-Enhanced Token Attribution (HETA): Interpreting Autoregressive LLMs

ICLR 2026poster

Attribution methods seek to explain language model predictions by quantifying the contribution of input tokens to generated outputs. However, most existing techniques are designed for encoder-based architectures and rely on linear approximations that fail to capture the causal and semantic complexit…

Cited by 1SourcecodeScholar
2026

Jailbreaking the Matrix: Nullspace Steering for Controlled Model Subversion

ICLR 2026poster

Large language models remain vulnerable to jailbreak attacks, inputs crafted to bypass safety mechanisms and elicit harmful responses, despite advances in alignment and instruction tuning. Existing attacks often rely on prompt rewrites, dense optimization, or ad hoc heuristics, and lack interpretabi…

Cited by 0SourcecodeScholar
2026

Selective Amnesia using Contrastive Subnet Erasure for Class Level Unlearning in Vision Models

CVPR 2026

We study concept-level forgetting in pretrained vision models: removing an entire semantic category so the system no longer recognizes that object in unseen images and contexts, rather than merely forgetting specific training examples. Prior work either applies blunt global projections or fine-tunes

Cited by 0SourcecodeScholar
2025

Grammar-Forced Translation of Natural Language to Temporal Logic using LLMs

ICML 2025poster

Translating natural language (NL) into a formal language such as temporal logic (TL) is integral for human communication with robots and autonomous systems. State-of-the-art approaches decompose the task into a grounding of atomic propositions (APs) phase and a translation phase. However, existing m…

Cited by 0SourcePDFScholar
2022

Shaping Noise for Robust Attributions in Neural Stochastic Differential Equations

AAAI 2022technical

Neural SDEs with Brownian motion as noise lead to smoother attributions than traditional ResNets. Various attribution methods such as saliency maps, integrated gradients, DeepSHAP and DeepLIFT have been shown to be more robust for neural SDEs than for ResNets using the recently proposed sensitivity…

Cited by 12SourcePDFScholar