← Search

Shaan Shah

2 accepted papers

2026

Representational Alignment Across Model Layers and Brain Regions with Hierarchical Optimal Transport

ICLR 2026poster

Standard representational similarity methods align each layer of a network to its best match in another independently, producing asymmetric results, lacking a global alignment score, and struggling with networks of different depths. These limitations arise from ignoring global activation structure a…

Cited by 0SourceScholar
2026

Safety Subspaces are Not Linearly Distinct: A Fine-Tuning Case Study

ICLR 2026poster

Large Language Models (LLMs) rely on safety alignment to produce socially acceptable responses. However, this behavior is known to be brittle: further fine-tuning, even on benign or lightly contaminated data, can degrade safety and reintroduce harmful behaviors. A growing body of work suggests that…

Cited by 0SourcecodeScholar