← Search

Boris Katz

22 accepted papers

2025

Population Transformer: Learning Population-level Representations of Neural Activity

ICLR 2025oral

We present a self-supervised framework that learns population-level codes for arbitrary ensembles of neural recordings at scale. We address key challenges in scaling models with neural time-series data, namely, sparse and variable electrode distribution across subjects and datasets. The Population T…

2025

Training the Untrainable: Introducing Inductive Bias via Representational Alignment

NeurIPS 2025poster

We demonstrate that architectures which traditionally are considered to be ill-suited for a task can be trained using inductive biases from another architecture. We call a network untrainable when it overfits, underfits, or converges to poor results even when tuning their hyperparameters. For examp…

Cited by 0SourceScholar
2024

Brain Treebank: Large-scale intracranial recordings from naturalistic language stimuli

NeurIPS 2024oral

We present the Brain Treebank, a large-scale dataset of electrophysiological neural responses, recorded from intracranial probes while 10 subjects watched one or more Hollywood movies. Subjects watched on average 2.6 Hollywood movies, for an average viewing time of 4.3 hours, and a total of 43 hours…

Cited by 3SourcePDFScholar
2024

BrainBits: How Much of the Brain are Generative Reconstruction Methods Using?

NeurIPS 2024poster

When evaluating stimuli reconstruction results it is tempting to assume that higher fidelity text and image generation is due to an improved understanding of the brain or more powerful signal extraction from neural recordings. However, in practice, new reconstruction methods could improve performan…

Cited by 0SourcePDFScholar
2024

Revealing Vision-Language Integration in the Brain with Multimodal Networks

ICML 2024poster

We use (multi)modal deep neural networks (DNNs) to probe for sites of multimodal integration in the human brain by predicting stereoencephalography (SEEG) recordings taken while human subjects watched movies. We operationalize sites of multimodal integration as regions where a multimodal vision-lang…

2023

BrainBERT: Self-supervised representation learning for intracranial recordings

ICLR 2023poster

We create a reusable Transformer, BrainBERT, for intracranial recordings bringing modern representation learning approaches to neuroscience. Much like in NLP and speech recognition, this Transformer enables classifying complex concepts, i.e., decoding neural data, with higher accuracy and with much…

2023

How hard are computer vision datasets? Calibrating dataset difficulty to viewing time

NeurIPS 2023poster

Humans outperform object recognizers despite the fact that models perform well on current datasets, including those explicitly designed to challenge machines with debiased images or distribution shift. This problem persists, in part, because we have no guidance on the absolute difficulty of an image…

Cited by 19SourcePDFScholar
2023

Zero-Shot Linear Combinations of Grounded Social Interactions with Linear Social MDPs

AAAI 2023technical

Humans and animals engage in rich social interactions. It is often theorized that a relatively small number of basic social interactions give rise to the full range of behavior observed. But no computational theory explaining how social interactions combine together has been proposed before. We do s…

Cited by 1SourcePDFScholar
2022

Incorporating Rich Social Interactions Into MDPs

ICRA 2022poster

Much of what we do as humans is engage socially with other agents, a skill that robots must also eventually possess. We demonstrate that a rich theory of social interactions originating from microsociology can be formalized by extending a nested MDP where agents reason about arbitrary functions of e…

Cited by 10SourceScholar
2022

The Aligned Multimodal Movie Treebank: An audio, video, dependency-parse treebank

EMNLP 2022main

Treebanks have traditionally included only text and were derived from written sources such as newspapers or the web. We introduce the Aligned Multimodal Movie Treebank (AMMT), an English language treebank derived from dialog in Hollywood movies which includes transcriptions of the audio-visual strea…

2022

Trajectory Prediction with Linguistic Representations

ICRA 2022poster

Language allows humans to build mental models that interpret what is happening around them resulting in more accurate long-term predictions. We present a novel trajectory prediction model that uses linguistic intermediate representations to forecast trajectories, and is trained using trajectory samp…

Cited by 22SourceScholar
2021

Compositional Networks Enable Systematic Generalization for Grounded Language Understanding

EMNLP 2021finding

Humans are remarkably flexible when understanding new sentences that include combinations of concepts they have never encountered before. Recent work has shown that while deep networks can mimic some human language abilities when presented with novel sentences, systematic variation uncovers the limi…

2021

Multi-resolution modeling of a discrete stochastic process identifies causes of cancer

ICLR 2021poster

Detection of cancer-causing mutations within the vast and mostly unexplored human genome is a major challenge. Doing so requires modeling the background mutation rate, a highly non-stationary stochastic process, across regions of interest varying in size from one to millions of positions. Here, we p…

Cited by 3SourcePDFScholar
2021

Neural Regression, Representational Similarity, Model Zoology & Neural Taskonomy at Scale in Rodent Visual Cortex

NeurIPS 2021poster

How well do deep neural networks fare as models of mouse visual cortex? A majority of research to date suggests results far more mixed than those produced in the modeling of primate visual cortex. Here, we perform a large-scale benchmarking of dozens of deep neural network models in mouse visual cor…

2021

PHASE: PHysically-grounded Abstract Social Events for Machine Social Perception

AAAI 2021technical

The ability to perceive and reason about social interactions in the context of physical environments is core to human social intelligence and human-machine cooperation. However, no prior dataset or benchmark has systematically evaluated physically grounded perception of complex social interactions t…

Cited by 36SourcePDFScholar
2020

Encoding formulas as deep networks: Reinforcement learning for zero-shot execution of LTL formulas

IROS 2020poster

We demonstrate a reinforcement learning agent which uses a compositional recurrent neural network that takes as input an LTL formula and determines satisfying actions. The input LTL formulas have never been seen before, yet the network performs zero-shot generalization to satisfy them. This is a nov…

Cited by 57SourcecodeScholar
2020

Learning a natural-language to LTL executable semantic parser for grounded robotics

CoRL 2020

Children acquire their native language with apparent ease by observing how language is used in context and attempting to use it themselves. They do so without laborious annotations, negative examples, or even direct corrections. We take a step toward robots that can do the same by training a grounde

Cited by 0SourcePDFScholar
2019

ObjectNet: A large-scale bias-controlled dataset for pushing the limits of object recognition models

NeurIPS 2019poster

We collect a large real-world test set, ObjectNet, for object recognition with controls where object backgrounds, rotations, and imaging viewpoints are random. Most scientific experiments have controls, confounds which are removed from the data, to ensure that subjects cannot perform a task by explo…

Cited by 697SourcePDFScholar