← Search

Lukas Cavigelli

5 accepted papers

2026

SINQ: Sinkhorn-Normalized Quantization for Calibration-Free Low-Precision LLM Weights

ICML 2026poster

Post-training quantization has emerged as the most widely used strategy for deploying large language models at low precision. Still, current methods show perplexity degradation at bit-widths $\leq 4$, partly because representing outliers causes precision issues in parameters that share the same scal…

Cited by 5SourceScholar
2026

TyphoonMLA: A Mixed Naive-Absorb MLA Kernel For Shared Prefix

ICLR 2026poster

Multi-Head Latent Attention (MLA) is a recent attention mechanism adopted in state-of-the-art LLMs such as DeepSeek-v3 and Kimi K2. Thanks to its novel formulation, MLA allows two functionally equivalent but computationally distinct kernel implementations: naive and absorb. While the naive kernels (…

Cited by 0SourceScholar
2025

AC-LoRA: (Almost) Training-Free Access Control Aware Multi-Modal LLMs

NeurIPS 2025poster

Corporate LLMs are gaining traction for efficient knowledge dissemination and management within organizations. However, as current LLMs are vulnerable to leaking sensitive information, it has proven difficult to apply them in settings where strict access control is necessary. To this end, we design…

Cited by 0SourceScholar
2023

RL-based Stateful Neural Adaptive Sampling and Denoising for Real-Time Path Tracing

NeurIPS 2023poster

Monte-Carlo path tracing is a powerful technique for realistic image synthesis but suffers from high levels of noise at low sample counts, limiting its use in real-time applications. To address this, we propose a framework with end-to-end training of a sampling importance network, a latent space enc…

2017

Soft-to-Hard Vector Quantization for End-to-End Learning Compressible Representations

NeurIPS 2017poster

We present a new approach to learn compressible representations in deep architectures with an end-to-end training strategy. Our method is based on a soft (continuous) relaxation of quantization and entropy, which we anneal to their discrete counterparts throughout training. We showcase this method…

Cited by 605SourcePDFScholar