← Search

Samuel Yeh

4 accepted papers

2026

LH-DECEPTION: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon Interactions

ICLR 2026poster

Deception is a pervasive feature of human communication and an emerging concern in large language models (LLMs). While recent studies document instances of LLM deception, most evaluations remain confined to single-turn prompts and fail to capture the long-horizon interactions in which deceptive stra…

Cited by 0SourceScholar
2026

LUMINA: Detecting Hallucinations in RAG System with Context–Knowledge Signals

ICLR 2026poster

Retrieval-Augmented Generation (RAG) aims to mitigate hallucinations in large language models (LLMs) by grounding responses in retrieved documents. Yet, RAG-based LLMs still hallucinate even when provided with correct and sufficient context. A growing line of work suggests that this stems from an im…

Cited by 0SourcecodeScholar
2025

Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment

NeurIPS 2025poster

Human feedback plays a pivotal role in aligning large language models (LLMs) with human preferences. However, such feedback is often noisy or inconsistent, which can degrade the quality of reward models and hinder alignment. While various automated data cleaning methods have been proposed to mitigat…

Cited by 0SourcecodeScholar
2025

MetaMind: Modeling Human Social Thoughts with Metacognitive Multi-Agent Systems

NeurIPS 2025spotlight

Human social interactions depend on the ability to infer others' unspoken intentions, emotions, and beliefs—a cognitive skill grounded in the psychological concept of Theory of Mind (ToM). While large language models (LLMs) excel in semantic understanding tasks, they struggle with the ambiguity and…

Cited by 0SourcecodeScholar