← Search

Manying Zhang

1 accepted papers

2025

On Weaponization-Resistant Large Language Models with Prospect Theoretic Alignment

COLING 2025main

Large language models (LLMs) have made significant advancements, but their increasing capabilities present serious risks of misuse, particularly in open-weight models where direct access to the model’s parameters is possible. Current safeguards, designed for closed-weight API models, are inadequate…

Cited by 1SourcePDFScholar