← Search

Khaoula Chehbouni

4 accepted papers

2025

Beyond the Safety Bundle: Auditing the Helpful and Harmless Dataset

NAACL 2025long

In an effort to mitigate the harms of large language models (LLMs), learning from human feedback (LHF) has been used to steer LLMs towards outputs that are intended to be both less harmful and more helpful. Despite the widespread adoption of LHF in practice, the quality of this feedback and its effe…

2025

Enhancing Privacy in the Early Detection of Sexual Predators Through Federated Learning and Differential Privacy

AAAI 2025technical

The increased screen time and isolation caused by the COVID-19 pandemic have led to a significant surge in cases of online grooming, which is the use of strategies by predators to lure children into sexual exploitation. Previous efforts to detect grooming in industry and academia have involved acces…

2025

Neither Valid nor Reliable? Investigating the Use of LLMs as Judges

NeurIPS 2025poster

Evaluating natural language generation (NLG) systems remains a core challenge, further complicated by the rise of general-purpose large language models (LLMs). Recently, large language models as judges (LLJs) have emerged as a scalable, cost-effective alternative to traditional metrics, but their va…

Cited by 0SourceScholar
2024

From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards

ACL 2024findings

Recent progress in large language models (LLMs) has led to their widespread adoption in various domains. However, these advancements have also introduced additional safety risks and raised concerns regarding their detrimental impact on already marginalized populations.Despite growing mitigation effo…