← Search

Renjie Gu

3 accepted papers

2025

MHALO: Evaluating MLLMs as Fine-grained Hallucination Detectors

ACL 2025finding

Hallucination remains a critical challenge for multimodal large language models (MLLMs), undermining their reliability in real-world applications. While fine-grained hallucination detection (FHD) holds promise for enhancing high-quality vision-language data construction and model alignment through e…

2025

Towards Resilient Safety-driven Unlearning for Diffusion Models against Downstream Fine-tuning

NeurIPS 2025poster

Text-to-image (T2I) diffusion models have achieved impressive image generation quality and are increasingly fine-tuned for personalized applications. However, these models often inherit unsafe behaviors from toxic pretraining data, raising growing safety concerns. While recent safety-driven unlearni…

Cited by 0SourcecodeScholar
2024

Course-Correction: Safety Alignment Using Synthetic Preferences

EMNLP 2024industry

The risk of harmful contents generated by large language models (LLMs) becomes a critical concern. This paper systematically evaluates and enhances LLMs’ capability to perform course-correction, , the model can steer away from generating harmful content autonomously. First, we introduce the C2-Eval…