← Search

Xudong Pan

9 accepted papers

2026

AutoControl Arena: Synthesizing Executable Test Environments for Frontier AI Risk Evaluation

ICML 2026poster

As Large Language Models (LLMs) evolve into autonomous agents, existing safety evaluations face a fundamental trade-off: manual benchmarks are costly, while LLM-based simulators are scalable but suffer from logic hallucination. We present AUTOCONTROL ARENA, an automated framework for frontier AI ris…

Cited by 0SourceScholar
2026

OpenDeception: Learning Deception and Trust in Human–AI Interaction via Multi-Agent Simulation

ICML 2026poster

As large language models (LLMs) are increasingly deployed as interactive agents, open-ended human-AI interactions can involve deceptive behaviors with serious real-world consequences, yet existing evaluations remain largely scenario-specific and model-centric. We introduce *OpenDeception*, a lightwe…

Cited by 0SourceScholar
2026

PRISON: Unmasking the Criminal Potential of Large Language Models

ICLR 2026poster

As large language models (LLMs) advance, concerns about their misconduct in complex social contexts intensify. Existing research has overlooked the systematic assessment of LLMs’ criminal potential in realistic interactions, where criminal potential is defined as the risk of producing harmful behavi…

Cited by 0SourceScholar
2026

Position: Preparing for AI Systems That Deceive Developers

ICML 2026poster

AI systems may exhibit deceptive behaviors that mislead developers about their capabilities, propensities, or actions. Such deception can take distinct forms across the development lifecycle: training subversion, evaluation gaming, and control evasion. We argue that the AI community should prioritiz…

Cited by 0SourceScholar
2026

Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction

ICML 2026poster

LLM-based agents solve complex tasks through iterative reasoning, tool use, and environment interaction, where each intermediate thought directly shapes subsequent actions. Small deviations in these thoughts can therefore propagate into unsafe behaviors, yet existing guardrails typically operate onl…

Cited by 0SourcecodeScholar
2023

RØROS: Building a Responsive Online Recommender System via Meta-Gradients Updating

ICASSP 2023accepted

In the era of information explosion, users of online services are urgently waiting for timely and effective recommendations. In this paper, we present the first study on the responsiveness aspect of recommender system and present Responsive Online RecOmmender System (RØROS) based on Meta-Gradients U…

Cited by 0SourceScholar
2022

House of Cans: Covert Transmission of Internal Datasets via Capacity-Aware Neuron Steganography

NeurIPS 2022accept

In this paper, we present a capacity-aware neuron steganography scheme (i.e., Cans) to covertly transmit multiple private machine learning (ML) datasets via a scheduled-to-publish deep neural network (DNN) as the carrier model. Unlike existing steganography schemes which treat the DNN parameters as…

Cited by 2SourcePDFScholar