← Search

Kai Rawal

2 accepted papers

2026

The Heterogeneous Safety Impacts of Benign Multilingual Fine-Tuning

ICML 2026poster

Fine-tuning a large language model is a ubiquitous method for enhancing its capability on a specific downstream task. However, prior work has shown that this increase in capability comes with a cost: it can increase a model's tendency to respond to unsafe adversarial prompts, even when fine-tuning w…

Cited by 0SourceScholar
2025

FairImagen: Post-Processing for Bias Mitigation in Text-to-Image Models

NeurIPS 2025poster

Text-to-image diffusion models, such as Stable Diffusion, have demonstrated remarkable capabilities in generating high-quality and diverse images from natural language prompts. However, recent studies reveal that these models often replicate and amplify societal biases, particularly along demographi…

Cited by 0SourcecodeScholar