← Search

Sina Mansouri

2 accepted papers

2026

Knowing Bias, Doing Better: Mitigating Social Bias in LLMs via Know-Bias Neuron Enhancement

ICML 2026poster

Large language models (LLMs) exhibit social biases that reinforce harmful stereotypes, limiting their safe deployment. Most existing debiasing methods adopt a suppressive paradigm by modifying parameters, prompts, or neurons associated with biased behavior; however, such approaches are often brittle…

Cited by 0SourceScholar
2026

Stability-Aware Feature Design for Robust Watermark Detection in Machine-Generated Text

ICML 2026poster

The widespread adoption of large language models (LLMs) has intensified the demand for principled methods to distinguish human- from machine-generated text. Watermarking provides a promising avenue, yet existing detectors exhibit sharp performance deterioration under multiple paraphrasing and when a…

Cited by 0SourceScholar