← Search

Minzhou Pan

6 accepted papers

2026

ARMs: Adaptive Red-Teaming Agent against Multimodal Models with Plug-and-Play Attacks

ICLR 2026poster

As vision-language models (VLMs) gain prominence, their multimodal interfaces also introduce new safety vulnerabilities, making the safety evaluation challenging and critical. Existing red-teaming efforts are either restricted to a narrow set of adversarial patterns or depend heavily on manual engin…

Cited by 0SourceScholar
2026

Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents

ICML 2026poster

LLM-based agents are becoming increasingly capable, yet their safety lags behind. This creates a gap between what agents can do and should do. This gap widens as agents engage in multi-turn interactions and employ diverse tools, introducing new risks overlooked by existing benchmarks. To systematica…

Cited by 0SourceScholar
2025

AIR-BENCH 2024: A Safety Benchmark based on Regulation and Policies Specified Risk Categories

ICLR 2025spotlight

Foundation models (FMs) provide societal benefits but also amplify risks. Governments, companies, and researchers have proposed regulatory frameworks, acceptable use policies, and safety benchmarks in response. However, existing public benchmarks often define safety categories based on previous lite…

Cited by 0SourcePDFScholar
2025

How to Evaluate and Mitigate IP Infringement in Visual Generative AI?

ICML 2025poster

The popularity of visual generative AI models like DALL-E 3, Stable Diffusion XL, Stable Video Diffusion, and Sora has been increasing. Through extensive evaluation, we discovered that the state-of-the-art visual generative models can generate content that bears a striking resemblance to characters…

Cited by 0SourcePDFScholar
2025

PolyGuard: Massive Multi-Domain Safety Policy-Grounded Guardrail Dataset

NeurIPS 2025poster

As large language models (LLMs) become widespread across diverse applications, concerns about the security and safety of LLM interactions have intensified. Numerous guardrail models and benchmarks have been developed to ensure LLM content safety. However, existing guardrail benchmarks are often buil…

Cited by 0SourceScholar
2025

SafeWatch: An Efficient Safety-Policy Following Video Guardrail Model with Transparent Explanations

ICLR 2025poster

With the rise of generative AI and rapid growth of high-quality video generation, video guardrails have become more crucial than ever to ensure safety and security across platforms. Current video guardrails, however, are either overly simplistic, relying on pure classification models trained on simp…