← Search

Qingchen Yu

4 accepted papers

2026

The Achilles’ Heel of LLMs: How Altering a Handful of Neurons Can Cripple Language Abilities

ICLR 2026poster

Large Language Models (LLMs) have become foundational tools in natural language processing, powering a wide range of applications and research. Many studies have shown that LLMs share significant similarities with the human brain. Neuroscience research has found that a small subset of biological neu…

Cited by 0SourcecodeScholar
2026

TurtleBench: Evaluating Top Language Models via Real-World Yes/No Puzzles

ICASSP 2026poster

As the application of Large Language Models (LLMs) expands, the demand for reliable evaluations increases. Existing LLM evaluation benchmarks primarily rely on static datasets, making it challenging to assess model performance in dynamic interactions with users. Moreover, these benchmarks often depe…

Cited by 0SourcePDFScholar
2025

GuessArena: Guess Who I Am? A Self-Adaptive Framework for Evaluating LLMs in Domain-Specific Knowledge and Reasoning

ACL 2025long

The evaluation of large language models (LLMs) has traditionally relied on static benchmarks, a paradigm that poses two major limitations: (1) predefined test sets lack adaptability to diverse application domains, and (2) standardized evaluation protocols often fail to capture fine-grained assessmen…

Cited by 0SourcePDFScholar
2025

xFinder: Large Language Models as Automated Evaluators for Reliable Evaluation

ICLR 2025poster

The continuous advancement of large language models (LLMs) has brought increasing attention to the critical issue of developing fair and reliable methods for evaluating their performance. Particularly, the emergence of cheating phenomena, such as test set leakage and prompt format overfitting, poses…

Cited by 0SourcePDFScholar