← Search

Emilio Castillo

2 accepted papers

2026

Sparser, Faster, Lighter Transformer Language Models

ICML 2026poster

Scaling autoregressive large language models (LLMs) has had an unprecedented impact, but at vast computational costs. In this work, we tackle these costs by leveraging unstructured sparsity within an LLM's feedforward layers, which account for the majority of its parameters and execution FLOPs. To a…

Cited by 0SourceScholar
2023

A fast heuristic to optimize time-space tradeoff for large models

NeurIPS 2023poster

Training large-scale neural networks is heavily constrained by GPU memory. In order to circumvent this limitation, gradient checkpointing, or recomputation is a powerful technique. There is active research in this area with methods such as Checkmake or Moccasin. However, both Checkmate and Moccasin…

Cited by 1SourcePDFScholar