← Search

Vivek Narayanaswamy

7 accepted papers

2026

Interpretable and Steerable Concept Bottleneck Sparse Autoencoders

CVPR 2026

Sparse autoencoders (SAEs) promise a unified approach for mechanistic interpretability, concept discovery, and model steering in LLMs and LVLMs. However, realizing this potential requires learned features to be both interpretable and steerable. To that end, we introduce two new computationally inexp

Cited by 0SourcecodeScholar
2024

On the Use of Anchoring for Training Vision Models

NeurIPS 2024spotlight

Anchoring is a recent, architecture-agnostic principle for training deep neural networks that has been shown to significantly improve uncertainty estimation, calibration, and extrapolation capabilities. In this paper, we systematically explore anchoring as a general protocol for training vision mode…

Cited by 0SourcePDFScholar
2024

PAGER: Accurate Failure Characterization in Deep Regression Models

ICML 2024poster

Safe deployment of AI models requires proactive detection of failures to prevent costly errors. To this end, we study the important problem of detecting failures in deep regression models. Existing approaches rely on epistemic uncertainty estimates or inconsistency w.r.t the training data to identif…

Cited by 3SourcePDFScholar
2022

Improved StyleGAN-v2 based Inversion for Out-of-Distribution Images

ICML 2022spotlight

Inverting an image onto the latent space of pre-trained generators, e.g., StyleGAN-v2, has emerged as a popular strategy to leverage strong image priors for ill-posed restoration. Several studies have showed that this approach is effective at inverting images similar to the data used for training. H…

2022

Single Model Uncertainty Estimation via Stochastic Data Centering

NeurIPS 2022accept

We are interested in estimating the uncertainties of deep neural networks, which play an important role in many scientific and engineering problems. In this paper, we present a striking new finding that an ensemble of neural networks with the same weight initialization, trained on datasets that are…

2021

Accurate and Robust Feature Importance Estimation under Distribution Shifts

AAAI 2021technical

With increasing reliance on the outcomes of black-box models in critical applications, post-hoc explainability tools that do not require access to the model internals are often used to enable humans understand and trust these models. In particular, we focus on the class of methods that can reveal th…

2021

Designing Counterfactual Generators using Deep Model Inversion

NeurIPS 2021poster

Explanation techniques that synthesize small, interpretable changes to a given image while producing desired changes in the model prediction have become popular for introspecting black-box models. Commonly referred to as counterfactuals, the synthesized explanations are required to contain discernib…

Cited by 27SourcePDFScholar