← Search

Peiran Dong

5 accepted papers

2025

Causally Motivated Sycophancy Mitigation for Large Language Models

ICLR 2025poster

Incorporating user preferences into large language models (LLMs) can enhance the personalization and reliability of model outputs and facilitate the application of LLMs to real-world scenarios. However, leveraging user preferences can be a double-edged sword. Recent studies have found that improper…

Cited by 0SourcePDFScholar
2024

Easing Concept Bleeding in Diffusion via Entity Localization and Anchoring

ICML 2024poster

Recent diffusion models have manifested extraordinary capabilities in generating high-quality, diverse, and innovative images guided by textual prompts. Nevertheless, these state-of-the-art models may encounter the challenge of concept bleeding when generating images with multiple entities or attrib…

Cited by 1SourcePDFScholar
2024

Towards Safe Concept Transfer of Multi-Modal Diffusion via Causal Representation Editing

NeurIPS 2024poster

Recent advancements in vision-language-to-image (VL2I) diffusion generation have made significant progress. While generating images from broad vision-language inputs holds promise, it also raises concerns about potential misuse, such as copying artistic styles without permission, which could have le…

Cited by 0SourcePDFScholar
2023

Towards Test-Time Refusals via Concept Negation

NeurIPS 2023poster

Generative models produce unbounded outputs, necessitating the use of refusal techniques to confine their output space. Employing generative refusals is crucial in upholding the ethical and copyright integrity of synthesized content, particularly when working with widely adopted diffusion models. "C…

Cited by 5SourcePDFScholar