← Search

Satvik Golechha

4 accepted papers

2026

Building Better Deception Probes Using Targeted Instruction Pairs

ICML 2026poster

Linear probes are a promising approach for monitoring AI systems for deceptive behaviour. Previous work has shown that a linear classifier trained on a contrastive instruction pair and a simple dataset can achieve good performance. However, these probes exhibit notable failures even in straightforwa…

Cited by 0SourceScholar
2025

A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders

NeurIPS 2025oral

Sparse Autoencoders (SAEs) aim to decompose the activation space of large language models (LLMs) into human-interpretable latent directions or features. As we increase the number of features in the SAE, hierarchical features tend to split into finer features (“math” may split into “algebra”, “geomet…

Cited by 0SourcecodeScholar
2024

NICE: To Optimize In-Context Examples or Not?

ACL 2024long

Recent work shows that in-context learning and optimization of in-context examples (ICE) can significantly improve the accuracy of large language models (LLMs) on a wide range of tasks, leading to an apparent consensus that ICE optimization is crucial for better performance. However, most of these s…