← Search

Grzegorz Gluch

6 accepted papers

2026

On the Impossibility of Separating Intelligence from Judgment: The Computational Intractability of Filtering for AI Alignment

ICLR 2026poster

With the increased deployment of large language models (LLMs), one concern is their potential misuse for generating harmful content. Our work studies the alignment challenge, with a focus on filters to prevent the generation of unsafe information. Two natural points of intervention are the filtering…

Cited by 0SourcecodeScholar
2025

The Good, the Bad and the Ugly: Meta-Analysis of Watermarks, Transferable Attacks and Adversarial Defenses

NeurIPS 2025poster

We formalize and analyze the trade-off between backdoor-based watermarks and adversarial defenses, framing it as an interactive protocol between a verifier and a prover. While previous works have primarily focused on this trade-off, our analysis extends it by identifying transferable attacks as a th…

Cited by 0SourceScholar
2023

Breaking a Classical Barrier for Classifying Arbitrary Test Examples in the Quantum Model

AISTATS 2023poster

A new model for adversarial robustness was introduced by Goldwasser et al. in [GKKM20]. In this model the authors present a selective and transductive learning algorithm which guarantees a low test error and low rejection rate wrt to the original distribution. Moreover, a lower bound in terms of the…

Cited by 2SourcePDFScholar
2020

Constructing a provably adversarially-robust classifier from a high accuracy one

AISTATS 2020poster

Modern machine learning models with very high accuracy have been shown to be vulnerable to small, adversarially chosen perturbations of the input. Given black-box access to a high-accuracy classifier f, we show how to construct a new classifier g that has high accuracy and is also robust to adversar…

Cited by 1SourcePDFScholar