← Search

Ahmed Elgohary

3 accepted papers

2025

Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements

ICLR 2025poster

The current paradigm for safety alignment of large language models (LLMs) follows a _one-size-fits-all_ approach: the model refuses to interact with any content deemed unsafe by the model provider. This approach lacks flexibility in the face of varying social norms across cultures and regions. In ad…

Cited by 0SourcePDFScholar
2025

Jailbreak Distillation: Renewable Safety Benchmarking

EMNLP 2025

Large language models (LLMs) are rapidly deployed in critical applications, raising urgent needs for robust safety benchmarking. We propose Jailbreak Distillation (JBDistill), a novel benchmark construction framework that “distills” jailbreak attacks into high-quality and easily-updatable safety ben

Cited by 0SourcePDFScholar
2021

NL-EDIT: Correcting Semantic Parse Errors through Natural Language Interaction

NAACL 2021long

We study semantic parsing in an interactive setting in which users correct errors with natural language feedback. We present NL-EDIT, a model for interpreting natural language feedback in the interaction context to generate a sequence of edits that can be applied to the initial parse to correct its…

Cited by 49SourcePDFScholar