← Search

Chenfu Bao

2 accepted papers

2026

Resisting Contextual Interference in RAG via Parametric-Knowledge Reinforcement

ICLR 2026poster

Retrieval-augmented generation (RAG) improves performance on knowledge-intensive tasks but can be derailed by wrong, irrelevant, or conflicting retrieved text, causing models to rely on inaccurate evidence and cascade errors. We propose Knowledgeable-R1, a reinforcement-learning framework that expli…

Cited by 0SourcecodeScholar
2025

Indirect Online Preference Optimization via Reinforcement Learning

IJCAI 2025

Human preference alignment (HPA) aims to ensure Large Language Models (LLMs) responding appropriately to meet human moral and ethical requirements. Existing methods, such as RLHF and DPO, rely heavily on high-quality human annotation, which restrict the efficiency of iterative online model refinemen

Cited by 0SourcePDFScholar