← Search

Panjia Qiu

2 accepted papers

2026

NDAD: Negative-Direction Aware Decoding for Large Language Models via Controllable Hallucination Signal Injection

ICLR 2026poster

Large language models (LLMs) have recently achieved impressive progress in knowledge-intensive and reasoning tasks. However, their tendency to produce fabricated or factually inconsistent content remains a fundamental challenge to their practical deployment. To address this issue, we propose Negativ…

Cited by 0SourceScholar
2025

LSSF: Safety Alignment for Large Language Models through Low-Rank Safety Subspace Fusion

ACL 2025long

The safety mechanisms of large language models (LLMs) exhibit notable fragility, as even fine-tuning on datasets without harmful content may still undermine their safety capabilities. Meanwhile, existing safety alignment methods predominantly rely on the fine-tuning process, which inadvertently lead…