← Search

Orestis Papakyriakopoulos

5 accepted papers

2026

Operationalizing Pluralistic Values in Large Language Model Alignment Reveals Trade-offs in Safety, Inclusivity, and Model Behavior

AAAI 2026technical

Although large language models (LLMs) are increasingly trained using human feedback for safety and alignment with human values, alignment decisions often overlook human social diversity. This study examines how incorporating pluralistic values affects LLM behavior by systematically evaluating demogr

Cited by 0SourcePDFScholar
2025

Information Retrieval Induced Safety Degradation in AI Agents

NeurIPS 2025poster

Despite the growing integration of retrieval-enabled AI agents into society, their safety and ethical behavior remain inadequately understood. In particular, the growing integration of LLMs and AI agents with external information sources and real-world environments raises critical questions about ho…

Cited by 0SourceScholar
2024

Position: Measure Dataset Diversity, Don't Just Claim It

ICML 2024oral

Machine learning (ML) datasets, often perceived as neutral, inherently encapsulate abstract and disputed social constructs. Dataset curators frequently employ value-laden terms such as diversity, bias, and quality to characterize datasets. Despite their prevalence, these terms lack clear definitions…

Cited by 17SourcePDFScholar
2024

Resampled Datasets Are Not Enough: Mitigating Societal Bias Beyond Single Attributes

EMNLP 2024main

We tackle societal bias in image-text datasets by removing spurious correlations between protected groups and image attributes. Traditional methods only target labeled attributes, ignoring biases from unlabeled ones. Using text-guided inpainting models, our approach ensures protected group independe…

Cited by 2SourcePDFScholar
2023

Ethical Considerations for Responsible Data Curation

NeurIPS 2023oral

Human-centric computer vision (HCCV) data curation practices often neglect privacy and bias concerns, leading to dataset retractions and unfair models. HCCV datasets constructed through nonconsensual web scraping lack crucial metadata for comprehensive fairness and robustness evaluations. Current re…