← Search

Usha Bhalla

3 accepted papers

2026

Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability

ICLR 2026oral

Translating the internal representations and computations of models into concepts that humans can understand is a key goal of interpretability. While recent dictionary learning methods such as Sparse Autoencoders (SAEs) provide a promising route to discover human-interpretable features, they often o…

Cited by 0SourcecodeScholar
2024

Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)

NeurIPS 2024poster

CLIP embeddings have demonstrated remarkable performance across a wide range of multimodal applications. However, these high-dimensional, dense vector representations are not easily interpretable, limiting our understanding of the rich structure of CLIP and its use in downstream applications that r…

2023

Discriminative Feature Attributions: Bridging Post Hoc Explainability and Inherent Interpretability

NeurIPS 2023poster

With the increased deployment of machine learning models in various real-world applications, researchers and practitioners alike have emphasized the need for explanations of model behaviour. To this end, two broad strategies have been outlined in prior literature to explain models. Post hoc explanat…