← Search

Bingshan Liu

2 accepted papers

2026

Value-Aligned Prompt Moderation via Zero-Shot Agentic Rewriting for Safe Image Generation

AAAI 2026technical

Generative vision-language models like Stable Diffusion demonstrate remarkable capabilities in creative media synthesis, but they also pose substantial risks of producing unsafe, offensive, or culturally inappropriate content when prompted adversarially. Current defenses struggle to align outputs wi

Cited by 0SourcePDFScholar
2025

Who Speaks for the Trigger? Dynamic Expert Routing in Backdoored Mixture-of-Experts Transformers

NeurIPS 2025poster

Large language models (LLMs) with Mixture-of-Experts (MoE) architectures achieve impressive performance and efficiency by dynamically routing inputs to specialized subnetworks, known as experts. However, this sparse routing mechanism inherently exhibits task preferences due to expert specialization…

Cited by 0SourceScholar