← Search

Linkang Yang

3 accepted papers

2026

Global-Local Confidence Fusion for Hallucination Detection in Mathematical Reasoning Task

AAAI 2026technical

Large Reasoning Models (LRMs) achieve promising results on complex reasoning tasks but remain susceptible to hallucinations. Existing hallucination detection methods based on Large Language Models (LLMs) often focus solely on final answers, overlooking inconsistencies between the answer and reasonin

Cited by 0SourcePDFScholar
2025

Dynamic Evil Score-Guided Decoding: An Efficient Decoding Framework For Red-Team Model

ACL 2025finding

Large language models (LLMs) have achieved significant advances but can potentially generate harmful content such as social biases, extremism, and misinformation. Red teaming is a promising approach to enhance model safety by creating adversarial prompts to test and improve model robustness. However…

Cited by 0SourcePDFScholar
2025

SafeConf: A Confidence-Calibrated Safety Self-Evaluation Method for Large Language Models

EMNLP 2025

Large language models (LLMs) have achieved groundbreaking progress in Natural Language Processing (NLP). Despite the numerous advantages of LLMs, they also pose significant safety risks. Self-evaluation mechanisms have gained increasing attention as a key safeguard to ensure safe and controllable co

Cited by 0SourcePDFScholar