← Search

Michael J. Tarr

13 accepted papers

2026

Meta-Learning In-Context Enables Training-Free Cross Subject Brain Decoding

CVPR 2026

Visual decoding from brain signals is a key challenge at the intersection of computer vision and neuroscience, requiring methods that bridge neural representations and computational models of vision. A field-wide goal is to achieve generalizable, cross-subject models. A major obstacle towards this g

Cited by 0SourcecodeScholar
2025

Brain Mapping with Dense Features: Grounding Cortical Semantic Selectivity in Natural Images With Vision Transformers

ICLR 2025poster

We introduce BrainSAIL (Semantic Attribution and Image Localization), a method for linking neural selectivity with spatially distributed semantic visual concepts in natural scenes. BrainSAIL leverages recent advances in large-scale artificial neural networks, using them to provide insights into the…

2025

Grounded Reinforcement Learning for Visual Reasoning

NeurIPS 2025poster

While reinforcement learning (RL) over chains of thought has significantly advanced language models in tasks such as mathematics and coding, visual reasoning introduces added complexity by requiring models to direct visual attention, interpret perceptual inputs, and ground abstract reasoning in spat…

Cited by 0SourcecodeScholar
2025

Meta-Learning an In-Context Transformer Model of Human Higher Visual Cortex

NeurIPS 2025poster

Understanding functional representations within higher visual cortex is a fundamental question in computational neuroscience. While artificial neural networks pretrained on large-scale datasets exhibit striking representational alignment with human neural responses, learning image-computable models…

Cited by 0SourceScholar
2025

Reanimating Images using Neural Representations of Dynamic Stimuli

CVPR 2025poster

While computer vision models have made incredible strides in static image recognition, they still do not match human performance in tasks that require the understanding of complex, dynamic motion. This is notably true for real-world scenarios where embodied agents face complex and motion-rich enviro…

2024

BrainSCUBA: Fine-Grained Natural Language Captions of Visual Cortex Selectivity

ICLR 2024poster

Understanding the functional organization of higher visual cortex is a central focus in neuroscience. Past studies have primarily mapped the visual and semantic selectivity of neural populations using hand-selected stimuli, which may potentially bias results towards pre-existing hypotheses of visual…

Cited by 12SourcePDFScholar
2024

Divergences between Language Models and Human Brains

NeurIPS 2024poster

Do machines and humans process language in similar ways? Recent research has hinted at the affirmative, showing that human neural activity can be effectively predicted using the internal representations of language models (LMs). Although such results are thought to reflect shared computational princ…

2024

VLM Agents Generate Their Own Memories: Distilling Experience into Embodied Programs of Thought

NeurIPS 2024spotlight

Large-scale generative language and vision-language models (LLMs and VLMs) excel in few-shot in-context learning for decision making and instruction following. However, they require high-quality exemplar demonstrations to be included in their context window. In this work, we ask: Can LLMs and VLMs g…

Cited by 5SourcePDFScholar
2023

Brain Diffusion for Visual Exploration: Cortical Discovery using Large Scale Generative Models

NeurIPS 2023oral

A long standing goal in neuroscience has been to elucidate the functional organization of the brain. Within higher visual cortex, functional accounts have remained relatively coarse, focusing on regions of interest (ROIs) and taking the form of selectivity for broad categories such as faces, places,…

Cited by 22SourcePDFScholar
2023

Brain Dissection: fMRI-trained Networks Reveal Spatial Selectivity in the Processing of Natural Images

NeurIPS 2023poster

The alignment between deep neural network (DNN) features and cortical responses currently provides the most accurate quantitative explanation for higher visual areas. At the same time, these model features have been critiqued as uninterpretable explanations, trading one black box (the human brain) f…

Cited by 9SourcePDFScholar
2023

Open-Ended Instructable Embodied Agents with Memory-Augmented Large Language Models

EMNLP 2023long findings

Pre-trained and frozen LLMs can effectively map simple scene re-arrangement instructions to programs over a robot's visuomotor functions through appropriate few-shot example prompting. To parse open-domain natural language and adapt to a user's idiosyncratic procedures, not known during prompt engin…

Cited by 0SourcecodeScholar
2022

Learning Neural Acoustic Fields

NeurIPS 2022accept

Our environment is filled with rich and dynamic acoustic information. When we walk into a cathedral, the reverberations as much as appearance inform us of the sanctuary's wide open space. Similarly, as an object moves around us, we expect the sound emitted to also exhibit this movement. While recent…

Cited by 79SourcePDFScholar
2022

TIDEE: Tidying Up Novel Rooms Using Visuo-Semantic Commonsense Priors

ECCV 2022poster

"We introduce TIDEE, an embodied agent that tidies up a disordered scene based on learned commonsense object placement and room arrangement priors. TIDEE explores a home environment, detects objects that are out of their natural place, infers plausible object contexts for them, localizes such contex…