← Search

Joshua Kazdan

9 accepted papers

2026

No, of Course I Can! Deeper Fine-Tuning Attacks That Bypass Token-Level Safety Mechanisms

ICLR 2026poster

Leading language model (LM) providers like OpenAI and Anthopic allow customers to fine-tune frontier LMs for specific use cases. To prevent abuse, these providers apply filters to block fine-tuning on overtly harmful data. In this setting, we make three contributions: First, while past work has show…

Cited by 0SourceScholar
2026

Position: Multiple Definitions & Unrealistic Assumptions of Model Collapse Distract from Real World Threats

ICML 2026poster

The proliferation of AI-generated content online has fueled concerns over \textit{model collapse}, a degradation in future generative models' performance when trained on synthetic data generated by earlier models. Industry leaders, premier research journals and popular science publications alike hav…

Cited by 0SourceScholar
2026

Truthfulness Does Not Scale Like Reasoning: Why Polling Fails as a Proxy Verifier

ICML 2026poster

Pass@$k$ and other methods of scaling inference compute can improve language model performance in domains with external verifiers, including mathematics and code, where incorrect candidates can be filtered reliably. This raises a natural question: can we similarly scale compute to elicit gains in tr…

Cited by 0SourceScholar
2025

CPSample: Classifier Protected Sampling for Guarding Training Data During Diffusion

ICLR 2025poster

Diffusion models have a tendency to exactly replicate their training data, especially when trained on small datasets. Most prior work has sought to mitigate this problem by imposing differential privacy constraints or masking parts of the training data, resulting in a notable substantial decrease i…

Cited by 2SourcePDFScholar
2025

Collapse or Thrive: Perils and Promises of Synthetic Data in a Self-Generating World

ICML 2025poster

What happens when generative machine learning models are pretrained on web-scale datasets containing data generated by earlier models? Some prior work warns of “model collapse” as the web is overwhelmed by synthetic data; other work suggests the problem can be contained (i.e. collapse can be avoided…

Cited by 8SourcePDFScholar
2025

How Do Large Language Monkeys Get Their Power (Laws)?

ICML 2025oral

Recent research across mathematical problem solving, proof assistant programming and multimodal jailbreaking documents a striking finding: when (multimodal) language model tackle a suite of tasks with multiple attempts per task -- succeeding if any attempt is correct -- then the negative log of the…

Cited by 0SourcePDFScholar
2025

KGGen: Extracting Knowledge Graphs from Plain Text with Language Models

NeurIPS 2025poster

Recent interest in building foundation models for knowledge graphs has highlighted a fundamental challenge: knowledge graph data is scarce. The best-known knowl- edge graphs are primarily human-labeled, created by pattern-matching, or extracted using early NLP techniques. While human-generated knowl…

Cited by 0SourcecodeScholar
2025

Position: Machine Learning Conferences Should Establish a "Refutations and Critiques" Track

NeurIPS 2025oral

Science progresses by iteratively advancing and correcting humanity's understanding of the world. In machine learning (ML) research, rapid advancements have led to an explosion of publications, but have also led to misleading, incorrect, flawed or perhaps even fraudulent studies being accepted and s…

Cited by 0SourceScholar
2025

The Utility and Complexity of In- and Out-of-Distribution Machine Unlearning

ICLR 2025poster

Machine unlearning, the process of selectively removing data from trained models, is increasingly crucial for addressing privacy concerns and knowledge gaps post-deployment. Despite this importance, existing approaches are often heuristic and lack formal guarantees. In this paper, we analyze the fun…

Cited by 1SourcePDFScholar