← Search

Yuxiao Lu

3 accepted papers

2026

Discern Truth from Falsehood: Reducing Over-Refusal via Contrastive Refinement

ICLR 2026poster

Large language models (LLMs) aligned for safety often suffer from over-refusal—the tendency to reject seemingly toxic or benign prompts by misclassifying them as toxic. This behavior undermines models' helpfulness and restricts usability in sensitive or nuanced contexts. While prior work has propose…

Cited by 0SourceScholar
2025

Semantic Loss Guided Data Efficient Supervised Fine Tuning for Safe Responses in LLMs

ICLR 2025poster

Large Language Models (LLMs) generating unsafe responses to toxic prompts is a significant issue in their applications. While various efforts aim to address this safety concern, previous approaches often demand substantial human data collection or rely on the less dependable option of using another…

Cited by 0SourcePDFScholar
2024

Handling Long and Richly Constrained Tasks through Constrained Hierarchical Reinforcement Learning

AAAI 2024technical

Safety in goal directed Reinforcement Learning (RL) settings has typically been handled through constraints over trajectories and have demonstrated good performance in primarily short horizon tasks. In this paper, we are specifically interested in the problem of solving temporally extended decision…

Cited by 0SourcePDFScholar