← Search

Abbas Rahimi

12 accepted papers

2026

A foundation model with multi-variate parallel attention to generate neuronal activity

ICLR 2026poster

Learning from multi-variate time-series with heterogeneous channel configurations remains a fundamental challenge for deep neural networks, particularly in clinical domains such as intracranial electroencephalography (iEEG), where channel setups vary widely across subjects. In this work, we introduc…

Cited by 0SourcecodeScholar
2026

Locally Coherent Parallel Decoding in Diffusion Language Models

ICML 2026poster

Diffusion language models (DLMs) have emerged as a promising alternative to autoregressive (AR) models, offering sub-linear generation latency and bidirectional capabilities that are particularly appealing for code generation and editing. Achieving sub-linear latency in discrete DLMs requires predic…

Cited by 0SourceScholar
2026

Thompson Sampling via Fine-Tuning of LLMs

ICLR 2026poster

Bayesian optimization in large unstructured discrete spaces is often hindered by the computational cost of maximizing acquisition functions due to the absence of gradients. We propose a scalable alternative based on Thompson sampling that eliminates the need for acquisition function maximization by…

Cited by 0SourcecodeScholar
2025

Analog Foundation Models

NeurIPS 2025poster

Analog in-memory computing (AIMC) is a promising compute paradigm to improve speed and power efficiency of neural network inference beyond the limits of conventional von Neumann-based architectures. However, AIMC introduces fundamental challenges such as noisy computations and strict constraints on…

Cited by 0SourcecodeScholar
2025

On the Expressiveness and Length Generalization of Selective State Space Models on Regular Languages

AAAI 2025technical

Selective state-space models (SSMs) are an emerging alternative to the Transformer, offering the unique advantage of parallel training and sequential inference. Although these models have shown promising performance on a variety of tasks, their formal expressiveness and length generalization propert…

2025

Scalable Evaluation and Neural Models for Compositional Generalization

NeurIPS 2025poster

Compositional generalization—a key open challenge in modern machine learning—requires models to predict unknown combinations of known concepts. However, assessing compositional generalization remains a fundamental challenge due to the lack of standardized evaluation protocols and the limitations of…

Cited by 0SourceScholar
2025

Structured Sparse Transition Matrices to Enable State Tracking in State-Space Models

NeurIPS 2025spotlight

Modern state-space models (SSMs) often utilize structured transition matrices which enable efficient computation but pose restrictions on the model’s expressivity, as measured in terms of the ability to emulate finite-state automata (FSA). While unstructured transition matrices are optimal in terms…

Cited by 0SourcecodeScholar
2025

The Case for Cleaner Biosignals: High-fidelity Neural Compressor Enables Transfer from Cleaner iEEG to Noisier EEG

ICLR 2025poster

All data modalities are not created equal, even when the signal they measure comes from the same source. In the case of the brain, two of the most important data modalities are the scalp electroencephalogram (EEG), and the intracranial electroencephalogram (iEEG). iEEG benefits from a higher signal-…

2024

Limits of Transformer Language Models on Learning to Compose Algorithms

NeurIPS 2024poster

We analyze the capabilities of Transformer language models in learning compositional discrete tasks. To this end, we evaluate training LLaMA models and prompting GPT-4 and Gemini on four tasks demanding to learn a composition of several discrete sub-tasks. In particular, we measure how well these mo…

2023

MIMONets: Multiple-Input-Multiple-Output Neural Networks Exploiting Computation in Superposition

NeurIPS 2023poster

With the advent of deep learning, progressively larger neural networks have been designed to solve complex tasks. We take advantage of these capacity-rich models to lower the cost of inference by exploiting computation in superposition. To reduce the computational burden per input, we propose Multip…

2022

Constrained Few-Shot Class-Incremental Learning

CVPR 2022poster

Continually learning new classes from fresh data without forgetting previous knowledge of old classes is a very challenging research problem. Moreover, it is imperative that such learning must respect certain memory and computational constraints such as (i) training samples are limited to only a few…

Cited by 187PDFcodeScholar