2026
Value-Aligned Prompt Moderation via Zero-Shot Agentic Rewriting for Safe Image Generation
AAAI 2026technical
Generative vision-language models like Stable Diffusion demonstrate remarkable capabilities in creative media synthesis, but they also pose substantial risks of producing unsafe, offensive, or culturally inappropriate content when prompted adversarially. Current defenses struggle to align outputs wi