← Search

Golnoosh Farnadi

18 accepted papers

2026

Neighbor-Aware Localized Concept Erasure in Text-to-Image Diffusion Models

CVPR 2026

Concept erasure in text-to-image diffusion models seeks to remove undesired concepts while preserving overall generative capability. Localized erasure methods aim to restrict edits to the spatial region occupied by the target concept. However, we observe that suppressing a concept can unintentionall

Cited by 0SourcecodeScholar
2025

Beyond the Safety Bundle: Auditing the Helpful and Harmless Dataset

NAACL 2025long

In an effort to mitigate the harms of large language models (LLMs), learning from human feedback (LHF) has been used to steer LLMs towards outputs that are intended to be both less harmful and more helpful. Despite the widespread adoption of LHF in practice, the quality of this feedback and its effe…

2025

Designing Ambiguity Sets for Distributionally Robust Optimization Using Structural Causal Optimal Transport

AAAI 2025technical

Distributionally robust optimization tackles out-of-sample issues like overfitting and distribution shifts by adopting an adversarial approach over a range of possible data distributions, known as the ambiguity set. To balance conservatism and accuracy, these sets must include realistic probability…

2025

Enhancing Privacy in the Early Detection of Sexual Predators Through Federated Learning and Differential Privacy

AAAI 2025technical

The increased screen time and isolation caused by the COVID-19 pandemic have led to a significant surge in cases of online grooming, which is the use of strategies by predators to lure children into sexual exploitation. Previous efforts to detect grooming in industry and academia have involved acces…

2025

Erasing More Than Intended? How Concept Erasure Degrades the Generation of Non-Target Concepts

ICCV 2025poster

Concept erasure techniques have recently gained significant attention for their potential to remove unwanted concepts from text-to-image models. While these methods often demonstrate promising results in controlled settings, their robustness in real-world applications and suitability for deployment…

Cited by 0SourcePDFScholar
2025

Hallucination Detox: Sensitivity Dropout (SenD) for Large Language Model Training

ACL 2025long

As large language models (LLMs) become increasingly prevalent, concerns about their reliability, particularly due to hallucinations - factually inaccurate or irrelevant outputs - have grown. Our research investigates the relationship between the uncertainty in training dynamics and the emergence of…

2025

Neither Valid nor Reliable? Investigating the Use of LLMs as Judges

NeurIPS 2025poster

Evaluating natural language generation (NLG) systems remains a core challenge, further complicated by the rise of general-purpose large language models (LLMs). Recently, large language models as judges (LLJs) have emerged as a scalable, cost-effective alternative to traditional metrics, but their va…

Cited by 33SourceScholar
2025

REVIVING YOUR MNEME: Predicting The Side Effects of LLM Unlearning and Fine-Tuning via Sparse Model Diffing

EMNLP 2025

LLMs are frequently fine-tuned or unlearned to adapt to new tasks or eliminate undesirable behaviors. While existing evaluation methods assess performance after such interventions, there remains no general approach for detecting unintended side effects—such as unlearning biology content degrading pe

Cited by 0SourcePDFScholar
2025

What Secrets Do Your Manifolds Hold? Understanding the Local Geometry of Generative Models

ICLR 2025poster

Deep Generative Models are frequently used to learn continuous representations of complex data distributions by training on a finite number of samples. For any generative model, including pre-trained foundation models with Diffusion or Transformer architectures, generation performance can significan…

2024

Balancing Act: Constraining Disparate Impact in Sparse Models

ICLR 2024poster

Model pruning is a popular approach to enable the deployment of large deep learning models on edge devices with restricted computational or storage capacities. Although sparse models achieve performance comparable to that of their dense counterparts at the level of the entire dataset, they exhibit h…

Cited by 4SourcePDFScholar
2024

Causal Adversarial Perturbations for Individual Fairness and Robustness in Heterogeneous Data Spaces

AAAI 2024technical

As responsible AI gains importance in machine learning algorithms, properties like fairness, adversarial robustness, and causality have received considerable attention in recent years. However, despite their individual significance, there remains a critical gap in simultaneously exploring and integr…

Cited by 4SourcePDFScholar
2024

From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards

ACL 2024findings

Recent progress in large language models (LLMs) has led to their widespread adoption in various domains. However, these advancements have also introduced additional safety risks and raised concerns regarding their detrimental impact on already marginalized populations.Despite growing mitigation effo…

2024

Learning to Build Solutions in Stochastic Matching Problems Using Flows (Student Abstract)

AAAI 2024technical

Generative Flow Networks, known as GFlowNets, have been introduced in recent times, presenting an exciting possibility for neural networks to model distributions across various data structures. In this paper, we broaden their applicability to encompass scenarios where the data structures are optimal…

Cited by 0SourcePDFScholar
2024

Position: Cracking the Code of Cascading Disparity Towards Marginalized Communities

ICML 2024poster

The rise of foundation models holds immense promise for advancing AI, but this progress may amplify existing risks and inequalities, leaving marginalized communities behind. In this position paper, we discuss that disparities towards marginalized communities – performance, representation, privacy, r…

Cited by 2SourcePDFScholar
2024

Promoting Fair Vaccination Strategies through Influence Maximization: A Case Study on COVID-19 Spread

AAAI 2024technical

The aftermath of the Covid-19 pandemic saw more severe outcomes for racial minority groups and economically-deprived communities. Such disparities can be explained by several factors, including unequal access to healthcare, as well as the inability of low income groups to reduce their mobility due t…

2024

Wasserstein Distributionally Robust Optimization through the Lens of Structural Causal Models and Individual Fairness

NeurIPS 2024poster

In recent years, Wasserstein Distributionally Robust Optimization (DRO) has garnered substantial interest for its efficacy in data-driven decision-making under distributional uncertainty. However, limited research has explored the application of DRO to address individual fairness concerns, particula…

Cited by 0SourcePDFScholar
2021

Individual Fairness in Kidney Exchange Programs

AAAI 2021technical

Kidney transplant is the preferred method of treatment for patients suffering from kidney failure. However, not all patients can find a donor which matches their physiological characteristics. Kidney exchange programs (KEPs) seek to match such incompatible patient-donor pairs together, usually with…

2020

Counterexample-Guided Learning of Monotonic Neural Networks

NeurIPS 2020poster

The widespread adoption of deep learning is often attributed to its automatic feature construction with minimal inductive bias. However, in many real-world tasks, the learned function is intended to satisfy domain-specific constraints. We focus on monotonicity constraints, which are common and requi…