← Search

Yuyan Bu

5 accepted papers

2026

Align Once, Benefit Multilingually: Enforcing Multilingual Consistency for LLM Safety Alignment

ICLR 2026poster

The widespread deployment of large language models (LLMs) across linguistic communities necessitates reliable multilingual safety alignment. However, recent efforts to extend alignment to other languages often require substantial resources, either through large-scale, high-quality supervision in the…

Cited by 0SourceScholar
2025

Beyond Excess and Deficiency: Adaptive Length Bias Mitigation in Reward Models for RLHF

NAACL 2025findings

Reinforcement Learning from Human Feedback (RLHF) is crucial for aligning large language models (LLMs) with human values. However, it has been noted that reward models in RLHF often exhibit unintended biases, such as an overemphasis on response length based on the erroneous assumption that longer re…

Cited by 0SourcePDFScholar
2025

The Staircase of Ethics: Probing LLM Value Priorities through Multi-Step Induction to Complex Moral Dilemmas

EMNLP 2025

Ethical decision-making is a critical aspect of human judgment, and the growing use of LLMs in decision-support systems necessitates a rigorous evaluation of their moral reasoning capabilities. However, existing assessments primarily rely on single-step evaluations, failing to capture how models ada

Cited by 0SourcePDFScholar
2023

FakeSV: A Multimodal Benchmark with Rich Social Context for Fake News Detection on Short Video Platforms

AAAI 2023technical

Short video platforms have become an important channel for news sharing, but also a new breeding ground for fake news. To mitigate this problem, research of fake news video detection has recently received a lot of attention. Existing works face two roadblocks: the scarcity of comprehensive and large…

2021

Self-Supervised Learning for Sleep Stage Classification with Predictive and Discriminative Contrastive Coding

ICASSP 2021accepted

The purpose of this paper is to learn efficient representations from raw electroencephalogram (EEG) signals for sleep stage classification via self-supervised learning (SSL). Although supervised methods have gained favorable performance, they heavily rely on manually labeled datasets. Recently, SSL…

Cited by 0SourceScholar