← Search

Maneesh Kumar Singh

3 accepted papers

2025

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain)

ICLR 2025poster

Transformer-based language models, though not explicitly trained to mimic brain recordings, have demonstrated surprising alignment with brain activity. Progress in these models—through increased size, instruction-tuning, and multimodality—has led to better representational alignment with neural data…

2025

Multi-modal brain encoding models for multi-modal stimuli

ICLR 2025poster

Despite participants engaging in unimodal stimuli, such as watching images or silent videos, recent work has demonstrated that multi-modal Transformer models can predict visual brain activity impressively well, even with incongruent modality representations. This raises the question of how accuratel…

2023

Unveiling The Mask of Position-Information Pattern Through the Mist of Image Features

ICML 2023poster

Recent studies have shown that paddings in convolutional neural networks encode absolute position information which can negatively affect the model performance for certain tasks. However, existing metrics for quantifying the strength of positional information remain unreliable and frequently lead to…

Cited by 3SourcePDFScholar