← Search

Juan Pablo Munoz

7 accepted papers

2025

Mamba-Shedder: Post-Transformer Compression for Efficient Selective Structured State Space Models

NAACL 2025long

Large pre-trained models have achieved outstanding results in sequence modeling. The Transformer block and its attention mechanism have been the main drivers of the success of these models. Recently, alternative architectures, such as Selective Structured State Space Models (SSMs), have been propose…

2024

EFTNAS: Searching for Efficient Language Models in First-Order Weight-Reordered Super-Networks

COLING 2024main

Transformer-based models have demonstrated outstanding performance in natural language processing (NLP) tasks and many other domains, e.g., computer vision. Depending on the size of these models, which have grown exponentially in the past few years, machine learning practitioners might be restricted…

2024

Federated Foundation Models: Privacy-Preserving and Collaborative Learning for Large Models

COLING 2024main

Foundation Models (FMs), such as LLaMA, BERT, GPT, ViT, and CLIP, have demonstrated remarkable success in a wide range of applications, driven by their ability to leverage vast amounts of data for pre-training. However, optimizing FMs often requires access to sensitive data, raising privacy concerns…

Cited by 59SourcePDFScholar
2024

LoNAS: Elastic Low-Rank Adapters for Efficient Large Language Models

COLING 2024main

Large Language Models (LLMs) continue to grow, reaching hundreds of billions of parameters and making it challenging for Deep Learning practitioners with resource-constrained systems to use them, e.g., fine-tuning these models for a downstream task of their interest. Adapters, such as low-rank adapt…

2024

SQFT: Low-cost Model Adaptation in Low-precision Sparse Foundation Models

EMNLP 2024finding

Large pre-trained models (LPMs), such as large language models, have become ubiquitous and are employed in many applications. These models are often adapted to a desired domain or downstream task through a fine-tuning stage. This paper proposes SQFT, an end-to-end solution for low-precision sparse p…

2022

EZNAS: Evolving Zero-Cost Proxies For Neural Architecture Scoring

NeurIPS 2022accept

Neural Architecture Search (NAS) has significantly improved productivity in the design and deployment of neural networks (NN). As NAS typically evaluates multiple models by training them partially or completely, the improved productivity comes at the cost of significant carbon footprint. To alleviat…

Cited by 14SourcePDFScholar