← Search

Yongkang Chen

4 accepted papers

2026

Beyond Mode Collapse: Distribution Matching for Diverse Reasoning

ICML 2026poster

On-policy reinforcement learning methods like GRPO suffer from \emph{mode collapse}: they exhibit reduced solution diversity, concentrating probability mass on a single solution once discovered and ceasing exploration of alternative strategies. We show this stems from reverse KL minimization's mode-…

Cited by 0SourceScholar
2025

Better Red Teaming via Searching with Large Language Model

ACL 2025finding

The safe deployment of large language models (LLMs) necessitates comprehensive safety evaluations through red teaming. However, existing methods face challenges in managing semantic intricacies and optimizing the efficiency of the search process. To overcome these limitations, we propose Better Red…

Cited by 1SourcePDFScholar
2025

Mixing Expert Knowledge: Bring Human Thoughts Back To the Game of Go

NeurIPS 2025poster

Large language models (LLMs) have demonstrated exceptional performance in reasoning tasks such as mathematics and coding, matching or surpassing human capabilities. However, these impressive reasoning abilities face significant challenges in specialized domains. Taking Go as an example, although Alp…

Cited by 0SourceScholar
2025

Refusal-Aware Red Teaming: Exposing Inconsistency in Safety Evaluations

EMNLP 2025

The responsible deployment of Large Language Models (LLMs) necessitates rigorous safety evaluations. However, a critical challenge arises from inconsistencies between an LLM’s internal refusal decisions and external safety assessments, hindering effective validation. This paper introduces the concep

Cited by 0SourcePDFScholar