← Search

Vardan Papyan

17 accepted papers

2025

Breach By A Thousand Leaks: Unsafe Information Leakage in 'Safe' AI Responses

ICLR 2025poster

Vulnerability of Frontier language models to misuse has prompted the development of safety measures like filters and alignment training seeking to ensure safety through robustness to adversarially crafted prompts. We assert that robustness is fundamentally insufficient for ensuring safety goals due…

Cited by 2SourcePDFScholar
2025

Transformer Block Coupling and its Correlation with Generalization in LLMs

ICLR 2025poster

Large Language Models (LLMs) have made significant strides in natural language processing, and a precise understanding of the internal mechanisms driving their success is essential. In this work, we analyze the trajectories of token embeddings as they pass through transformer blocks, linearizing the…

Cited by 0SourcePDFScholar
2024

Out of the Ordinary: Spectrally Adapting Regression for Covariate Shift

ICML 2024poster

Designing deep neural network classifiers that perform robustly on distributions differing from the available training data is an active area of machine learning research. However, out-of-distribution generalization for regression---the analogous problem for modeling continuous targets---remains rel…

Cited by 1SourcePDFScholar
2024

Position: Fundamental Limitations of LLM Censorship Necessitate New Approaches

ICML 2024poster

Large language models (LLMs) have exhibited impressive capabilities in comprehending complex instructions. However, their blind adherence to provided instructions has led to concerns regarding risks of malicious use. Existing defence mechanisms, such as model fine-tuning or output censorship methods…

Cited by 2SourcePDFScholar
2024

Sparsest Models Elude Pruning: An Exposé of Pruning’s Current Capabilities

ICML 2024poster

Pruning has emerged as a promising approach for compressing large-scale models, yet its effectiveness in recovering the sparsest of models has not yet been explored. We conducted an extensive series of 485,838 experiments, applying a range of state-of-the-art pruning algorithms to a synthetic datase…

2022

Neural Collapse Under MSE Loss: Proximity to and Dynamics on the Central Path

ICLR 2022oral

The recently discovered Neural Collapse (NC) phenomenon occurs pervasively in today's deep net training paradigm of driving cross-entropy (CE) loss towards zero. During NC, last-layer features collapse to their class-means, both classifiers and class-means collapse to the same Simplex Equiangular Ti…

2019

Measurements of Three-Level Hierarchical Structure in the Outliers in the Spectrum of Deepnet Hessians

ICML 2019oral

We expose a structure in deep classifying neural networks in the derivative of the logits with respect to the parameters of the model, which is used to explain the existence of outliers in the spectrum of the Hessian. Previous works decomposed the Hessian into two components, attributing the outlier…

Cited by 87SourcePDFScholar
2018

Neural Proximal Gradient Descent for Compressive Imaging

NeurIPS 2018poster

Recovering high-resolution images from limited sensory data typically leads to a serious ill-posed inverse problem, demanding inversion algorithms that effectively capture the prior information. Learning a good inverse mapping from training data faces severe challenges, including: (i) scarcity of tr…

2018

Projecting on to the Multi-Layer Convolutional Sparse Coding Model

ICASSP 2018accepted

The recently proposed Multi-Layer Convolutional Sparse Coding (ML-CSC) model, consisting of a cascade of convolutional sparse layers, provides a new interpretation of Convolutional Neural Networks (CNNs). Under this framework, the forward pass in a CNN is equivalent to an algorithm that estimates ne…

Cited by 0SourceScholar