← Search

Jayneel Parekh

8 accepted papers

2025

Analyzing Finetuning Representation Shift for Multimodal LLMs Steering

ICCV 2025poster

Multimodal LLMs (MLLMs) have reached remarkable levels of proficiency in understanding multimodal inputs. However, understanding and interpreting the behavior of such complex models is a challenging task, not to mention the dynamic shifts that may occur during fine-tuning, or due to covariate shift…

Cited by 0SourcePDFScholar
2025

Learning to Steer: Input-dependent Steering for Multimodal LLMs

NeurIPS 2025poster

Steering has emerged as a practical approach to enable post-hoc guidance of LLMs towards enforcing a specific behavior. However, it remains largely underexplored for multimodal LLMs (MLLMs); furthermore, existing steering techniques, such as \textit{mean} steering, rely on a single steering vector,…

Cited by 0SourceScholar
2025

One Wave To Explain Them All: A Unifying Perspective On Feature Attribution

ICML 2025poster

Feature attribution methods aim to improve the transparency of deep neural networks by identifying the input features that influence a model's decision. Pixel-based heatmaps have become the standard for attributing features to high-dimensional inputs, such as images, audio representations, and volum…

2025

Restyling Unsupervised Concept Based Interpretable Networks with Generative Models

ICLR 2025poster

Developing inherently interpretable models for prediction has gained prominence in recent years. A subclass of these models, wherein the interpretable network relies on learning high-level concepts, are valued because of closeness of concept representations to human communication. However, the visua…

2024

A Concept-Based Explainability Framework for Large Multimodal Models

NeurIPS 2024poster

Large multimodal models (LMMs) combine unimodal encoders and large language models (LLMs) to perform multimodal tasks. Despite recent advancements towards the interpretability of these models, understanding internal representations of LMMs remains largely a mystery. In this paper, we present a novel…

2022

Listen to Interpret: Post-hoc Interpretability for Audio Networks with NMF

NeurIPS 2022accept

This paper tackles post-hoc interpretability for audio processing networks. Our goal is to interpret decisions of a trained network in terms of high-level audio objects that are also listenable for the end-user. To this end, we propose a novel interpreter design that incorporates non-negative matrix…