← Search

Pinlong Zhao

5 accepted papers

2026

Aligner, Diagnose Thyself: A Meta-Learning Paradigm for Fusing Intrinsic Feedback in Preference Alignment

ICLR 2026poster

The alignment of Large Language Models (LLMs) with human preferences is critically undermined by noisy labels in training datasets. Existing robust methods often prove insufficient, as they rely on single, narrow heuristics such as perplexity or loss, failing to address the diverse nature of real-wo…

Cited by 0SourceScholar
2026

Robust LLM Unlearning via Post Judgment and Multi-round Thinking

ICLR 2026poster

The unlearning capability of LLMs is vital for ensuring compliance and safety, especially when removing sensitive knowledge from deployed models. Pre-filtering methods, enabling rapid deployment without parameter changes, are a prominent unlearning approach. However, they exhibit significant robustn…

Cited by 0SourcecodeScholar