← Search

Saleh Momeni

5 accepted papers

2026

Push, Pop, Parallelize: Stack-Augmented Linear Attention via the Delta Rule

ICML 2026poster

Linear attention architectures based on the Delta rule, such as DeltaNet and RWKV-7, combine Transformers' training scalability with RNNs' inference efficiency and can provably solve regular language tasks. However, due to their fixed-size state, these models fundamentally struggle to capture the re…

Cited by 0SourceScholar
2025

AnaCP: Toward Upper-Bound Continual Learning via Analytic Contrastive Projection

NeurIPS 2025spotlight

This paper studies the problem of class-incremental learning (CIL), a core setting within continual learning where a model learns a sequence of tasks, each containing a distinct set of classes. Traditional CIL methods, which do not leverage pre-trained models (PTMs), suffer from catastrophic forgett…

Cited by 0SourceScholar
2025

Continual Learning Using a Kernel-Based Method Over Foundation Models

AAAI 2025technical

Continual learning (CL) learns a sequence of tasks incrementally. This paper studies the challenging CL setting of class-incremental learning (CIL). CIL has two key challenges: catastrophic forgetting (CF) and inter-task class separation (ICS). Despite numerous proposed methods, these issues remain…

2025

In-context Continual Learning Assisted by an External Continual Learner

COLING 2025main

Existing continual learning (CL) methods mainly rely on fine-tuning or adapting large language models (LLMs). They still suffer from catastrophic forgetting (CF). Little work has been done to exploit in-context learning (ICL) to leverage the extensive knowledge within LLMs for CL without updating an…

Cited by 0SourcePDFScholar