← Search

Claudio Mayrink Verdun

12 accepted papers

2026

GradPCA: Leveraging NTK Alignment for Reliable Out-of-Distribution Detection

ICLR 2026poster

We introduce GradPCA, an Out-of-Distribution (OOD) detection method that exploits the low-rank structure of neural network gradients induced by Neural Tangent Kernel (NTK) alignment. GradPCA applies Principal Component Analysis (PCA) to gradient class-means, achieving more consistent performance tha…

Cited by 0SourcecodeScholar
2026

Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability

ICLR 2026oral

Translating the internal representations and computations of models into concepts that humans can understand is a key goal of interpretability. While recent dictionary learning methods such as Sparse Autoencoders (SAEs) provide a promising route to discover human-interpretable features, they often o…

Cited by 0SourcecodeScholar
2025

Get rid of your constraints and reparametrize: A study in NNLS and implicit bias

AISTATS 2025poster

Over the past years, there has been significant interest in understanding the implicit bias of gradient descent optimization and its connection to the generalization properties of overparametrized neural networks. Several works observed that when training linear diagonal networks on the square loss…

Cited by 0SourceScholar
2025

HeavyWater and SimplexWater: Distortion-free LLM Watermarks for Low-Entropy Distributions

NeurIPS 2025poster

Large language model (LLM) watermarks enable authentication of text provenance, curb misuse of machine-generated text, and promote trust in AI systems. Current watermarks operate by changing the next-token predictions output by an LLM. The updated (i.e., watermarked) predictions depend on random si…

Cited by 0SourceScholar
2025

Inference-Time Reward Hacking in Large Language Models

NeurIPS 2025spotlight

A common paradigm to improve the performance of large language models is optimizing for a reward model. Reward models assign a numerical score to an LLM’s output that indicates, for example, how likely it is to align with user preferences or safety goals. However, reward models are never perfect. Th…

Cited by 0SourceScholar
2025

Multi-Group Proportional Representations for Text-to-Image Models

CVPR 2025poster

Text-to-image (T2I) generative models can create vivid, realistic images from textual descriptions. As these models proliferate, they expose new concerns about their ability to represent diverse demographic groups, propagate stereotypes, and efface minority populations. Despite growing attention to…

2024

Imaging with Confidence: Uncertainty Quantification for High-dimensional Undersampled MR Images

ECCV 2024poster

"Establishing certified uncertainty quantification (UQ) in imaging processing applications continues to pose a significant challenge. In particular, such a goal is crucial for accurate and reliable medical imaging if one aims for precise diagnostics and appropriate intervention. In the case of magne…

2024

Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models

NeurIPS 2024poster

What latent features are encoded in language model (LM) representations? Recent work on training sparse autoencoders (SAEs) to disentangle interpretable features in LM representations has shown significant promise. However, evaluating the quality of these SAEs is difficult because we lack a ground-t…

2024

Multi-Group Proportional Representation in Retrieval

NeurIPS 2024poster

Image search and retrieval tasks can perpetuate harmful stereotypes, erase cultural identities, and amplify social disparities. Current approaches to mitigate these representational harms balance the number of retrieved items across population groups defined by a small number of (often binary) attri…

2024

Non-Asymptotic Uncertainty Quantification in High-Dimensional Learning

NeurIPS 2024spotlight

Uncertainty quantification (UQ) is a crucial but challenging task in many high-dimensional learning problems to increase the confidence of a given predictor. We develop a new data-driven approach for UQ in regression that applies both to classical optimization approaches such as the LASSO as well as…

2023

High-Dimensional Confidence Regions in Sparse MRI

ICASSP 2023accepted

One of the most promising solutions for uncertainty quantification in high-dimensional statistics is the debiased LASSO that relies on unconstrained ℓ <inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</inf> -minimization. The initial works focused on re…

Cited by 0SourceScholar
2021

Iteratively Reweighted Least Squares for Basis Pursuit with Global Linear Convergence Rate

NeurIPS 2021spotlight

The recovery of sparse data is at the core of many applications in machine learning and signal processing. While such problems can be tackled using $\ell_1$-regularization as in the LASSO estimator and in the Basis Pursuit approach, specialized algorithms are typically required to solve the correspo…

Cited by 22SourcePDFScholar