← Search

Jingxuan Fan

4 accepted papers

2026

ENCORE: Entropy-guided Reward Composition for Multi-head Safety Reward Models

AAAI 2026technical

The safety alignment of large language models (LLMs) often relies on reinforcement learning from human feedback (RLHF), which requires human annotations to construct preference datasets. Given the challenge of assigning overall quality scores to data, recent works increasingly adopt fine-grained rat

Cited by 0SourcePDFScholar
2025

HARDMath: A Benchmark Dataset for Challenging Problems in Applied Mathematics

ICLR 2025poster

Advanced applied mathematics problems are underrepresented in existing Large Language Model (LLM) benchmark datasets. To address this, we introduce $\textbf{HARDMath}$, a dataset inspired by a graduate course on asymptotic methods, featuring challenging applied mathematics problems that require anal…

2025

RuleAdapter: Dynamic Rules for training Safety Reward Models in RLHF

ICML 2025poster

Reinforcement Learning from Human Feedback (RLHF) is widely used to align models with human preferences, particularly to enhance the safety of responses generated by LLMs. This method traditionally relies on choosing preferred responses from response pairs. However, due to variations in human opinio…

Cited by 0SourcePDFScholar