← Search

Abhinav Sukumar Rao

3 accepted papers

2025

NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models

NAACL 2025long

To be effectively and safely deployed to global user populations, large language models (LLMs) may need to adapt outputs to user values and cultures, not just know about them. We introduce NormAd, an evaluation framework to assess LLMs’ cultural adaptability, specifically measuring their ability to…

2024

Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks

COLING 2024main

Recent explorations with commercial Large Language Models (LLMs) have shown that non-expert users can jailbreak LLMs by simply manipulating their prompts; resulting in degenerate output behavior, privacy and security breaches, offensive outputs, and violations of content regulator policies. Limited…

2023

Ethical Reasoning over Moral Alignment: A Case and Framework for In-Context Ethical Policies in LLMs

EMNLP 2023long findings

In this position paper, we argue that instead of morally aligning LLMs to specific set of ethical principles, we should infuse generic ethical reasoning capabilities into them so that they can handle value pluralism at a global scale. When provided with an ethical policy, an LLM should be capable of…

Cited by 0SourceScholar