← Search

Xinnuo Li

2 accepted papers

2026

Sycophancy Towards Researchers Drives Performative Misalignment

ICML 2026spotlight

The increasing situational awareness of language models raises safety concerns: models might be aware when they are evaluated, and adjust their behavior to evade monitoring and resist modification, e.g., pretending to be aligned only in evaluation. This \emph{alignment faking} behavior is often inte…

Cited by 0SourceScholar
2025

Revolve: Optimizing AI Systems by Tracking Response Evolution in Textual Optimization

ICML 2025poster

Recent advancements in large language models (LLMs) have significantly enhanced the ability of LLM-based systems to perform complex tasks through natural language processing and tool interaction. However, optimizing these LLM-based systems for specific tasks remains challenging, often requiring manu…