← Search

Pratyush Chatterjee

2 accepted papers

2026

AURA: Affordance-Understanding and Risk-aware Alignment Technique for Large Language Models

AAAI 2026technical

Present day LLMs face the challenge of managing affordance-based safety risks—situations where outputs inadvertently facilitate harmful actions due to overlooked logical implications. Traditional safety solutions, such as scalar outcome-based reward models, parameter tuning, or heuristic decoding st

Cited by 0SourcePDFScholar
2025

Soteria: Language-Specific Functional Parameter Steering for Multilingual Safety Alignment

EMNLP 2025

Ensuring consistent safety across multiple languages remains a significant challenge for large language models (LLMs). We introduce Soteria, a lightweight yet powerful strategy that locates and minimally adjusts the “functional heads” most responsible for harmful content generation in each language.