← Search

Noveen Sachdeva

4 accepted papers

2026

How to train data-efficient LLMs

ICLR 2026poster

The training of large language models (LLMs) is expensive. In this paper, we study data-efficient approaches for pre-training LLMs, \ie, techniques that aim to optimize the Pareto frontier of model quality and training resource/data consumption. We seek to understand the tradeoffs associated with da…

Cited by 0SourceScholar
2025

ActionPiece: Contextually Tokenizing Action Sequences for Generative Recommendation

ICML 2025spotlight

Generative recommendation (GR) is an emerging paradigm where user actions are tokenized into discrete token patterns and autoregressively generated as predictions. However, existing GR models tokenize each action independently, assigning the same fixed tokens to identical actions across all sequence…

2025

GC4NC: A Benchmark Framework for Graph Condensation on Node Classification with New Insights

NeurIPS 2025poster

Graph condensation (GC) is an emerging technique designed to learn a significantly smaller graph that retains the essential information of the original graph. This condensed graph has shown promise in accelerating graph neural networks while preserving performance comparable to those achieved with t…

Cited by 0SourcecodeScholar
2022

Infinite Recommendation Networks: A Data-Centric Approach

NeurIPS 2022accept

We leverage the Neural Tangent Kernel and its equivalence to training infinitely-wide neural networks to devise $\infty$-AE: an autoencoder with infinitely-wide bottleneck layers. The outcome is a highly expressive yet simplistic recommendation model with a single hyper-parameter and a closed-form s…

Cited by 30SourcePDFScholar