← Search

Yahan Yang

3 accepted papers

2025

MrGuard: A Multilingual Reasoning Guardrail for Universal LLM Safety

EMNLP 2025

Large Language Models (LLMs) are susceptible to adversarial attacks such as jailbreaking, which can elicit harmful or unsafe behaviors. This vulnerability is exacerbated in multilingual settings, where multilingual safety-aligned data is often limited. Thus, developing a guardrail capable of detecti

2023

Bootstrapping Small \& High Performance Language Models with Unmasking-Removal Training Policy

EMNLP 2023short main

BabyBERTa, a language model trained on small-scale child-directed speech while none of the words are unmasked during training, has been shown to achieve a level of grammaticality comparable to that of RoBERTa-base, which is trained on 6,000 times more words and 15 times more parameters. Relying on t…

Cited by 0SourceScholar