← Search

Aritra Das

2 accepted papers

2024

Learning to grok: Emergence of in-context learning and skill composition in modular arithmetic tasks

NeurIPS 2024oral

Large language models can solve tasks that were not present in the training set. This capability is believed to be due to in-context learning and skill composition. In this work, we study the emergence of in-context learning and skill composition in a collection of modular arithmetic tasks. Specific…

2024

To Grok or not to Grok: Disentangling Generalization and Memorization on Corrupted Algorithmic Datasets

ICLR 2024poster

Robust generalization is a major challenge in deep learning, particularly when the number of trainable parameters is very large. In general, it is very difficult to know if the network has memorized a particular set of examples or understood the underlying rule (or both). Motivated by this challenge…