← Search

Mayur Srungarapu

1 accepted papers

2026

Configurable Reward Model for Balanced Safety Alignment

ICML 2026poster

Aligning large language models (LLMs) to heterogeneous and rapidly evolving safety requirements remains a critical challenge. Existing instruction-tuned LLMs and standalone safety classifiers often fail to generalize to new safety configurations, motivating the need for Reward Models (RMs) that are …

Cited by 0SourceScholar