← Search

Jonathan Kahana

6 accepted papers

2025

Deep Linear Probe Generators for Weight Space Learning

ICLR 2025poster

Weight space learning aims to extract information about a neural network, such as its training dataset or generalization error. Recent approaches learn directly from model weights, but this presents many challenges as weights are high-dimensional and include permutation symmetries between neurons. A…

Cited by 2SourcePDFScholar
2025

We Should Chart an Atlas of All the World's Models

NeurIPS 2025poster

Public model repositories now contain millions of models, yet most remain undocumented and effectively lost: their capabilities, provenance, and constraints cannot be reliably determined. As a result, the field wastes training time and compute, propagates hidden biases, faces intellectual-property r…

Cited by 0SourceScholar
2024

Recovering the Pre-Fine-Tuning Weights of Generative Models

ICML 2024poster

The dominant paradigm in generative modeling consists of two steps: i) pre-training on a large-scale but unsafe dataset, ii) aligning the pre-trained model with human values via fine-tuning. This practice is considered safe, as no current method can recover the unsafe, *pre-fine-tuning* model weight…

2023

Red PANDA: Disambiguating Image Anomaly Detection by Removing Nuisance Factors

ICLR 2023poster

Anomaly detection methods strive to discover patterns that differ from the norm in a meaningful way. This goal is ambiguous as different human operators may find different attributes meaningful. An image differing from the norm by an attribute such as pose may be considered anomalous by some operato…

Cited by 4SourcePDFScholar