← Search

Daniele Paliotta

4 accepted papers

2024

The Mamba in the Llama: Distilling and Accelerating Hybrid Models

NeurIPS 2024poster

Linear RNN architectures, like Mamba, can be competitive with Transformer models in language modeling while having advantageous deployment characteristics. Given the focus on training large-scale Transformer models, we consider the challenge of converting these pretrained models for deployment. We…

2024

Understanding and Minimising Outlier Features in Transformer Training

NeurIPS 2024poster

Outlier Features (OFs) are neurons whose activation magnitudes significantly exceed the average over a neural network's (NN) width. They are well known to emerge during standard transformer training and have the undesirable effect of hindering quantisation in afflicted models. Despite their practica…

Cited by 2SourcePDFScholar
2023

Fast Attention Over Long Sequences With Dynamic Sparse Flash Attention

NeurIPS 2023poster

Transformer-based language models have found many diverse applications requiring them to process sequences of increasing length. For these applications, the causal self-attention---which is the only component scaling quadratically w.r.t. the sequence length---becomes a central concern. While many wo…

Cited by 12SourcePDFScholar
2023

SUPA: A Lightweight Diagnostic Simulator for Machine Learning in Particle Physics

NeurIPS 2023poster

Deep learning methods have gained popularity in high energy physics for fast modeling of particle showers in detectors. Detailed simulation frameworks such as the gold standard \textsc{Geant4} are computationally intensive, and current deep generative architectures work on discretized, lower resolut…

Cited by 3SourcePDFScholar