← Search

Guy N. Rothblum

6 accepted papers

2026

On the Impossibility of Separating Intelligence from Judgment: The Computational Intractability of Filtering for AI Alignment

ICLR 2026poster

With the increased deployment of large language models (LLMs), one concern is their potential misuse for generating harmful content. Our work studies the alignment challenge, with a focus on filters to prevent the generation of unsafe information. Two natural points of intervention are the filtering…

Cited by 0SourcecodeScholar
2025

A Theory for Worst-Case vs. Average-Case Guarantees for LLMs

NeurIPS 2025poster

How can we trust the correctness of a learned model on a particular input of interest? Model accuracy is typically measured *on average* over a distribution of inputs, giving no guarantee for any fixed input. This paper proposes a theoretically-founded solution to this problem: to train *Self-Provin…

Cited by 0SourceScholar
2025

How to Verify Any (Reasonable) Distribution Property: Computationally Sound Argument Systems for Distributions

ICLR 2025poster

As statistical analyses become more central to science, industry and society, there is a growing need to ensure correctness of their results. Approximate correctness can be verified by replicating the entire analysis, but can we verify without replication? We focus on distribution testing problems:…

Cited by 0SourcePDFScholar
2025

PREAMBLE: Private and Efficient Aggregation via Block Sparse Vectors

NeurIPS 2025poster

We revisit the problem of secure aggregation of high-dimensional vectors in a two-server system such as Prio. These systems are typically used to aggregate vectors such as gradients in private federated learning, where the aggregate itself is protected via noise addition to ensure differential priva…

Cited by 0SourceScholar