← Search

Zhuohan Long

3 accepted papers

2025

Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

COLING 2025main

This paper presents a benchmark self-evolving framework to dynamically evaluate rapidly advancing Large Language Models (LLMs). We utilize a multi-agent system to reframe new evolving instances with high confidence that extend existing benchmarks. Towards a more scalable, robust and fine-grained eva…

2025

How Jailbreak Defenses Work and Ensemble? A Mechanistic Investigation

EMNLP 2025

Jailbreak attacks, where harmful prompts bypass generative models’ built-in safety, raise serious concerns about model vulnerability. While many defense methods have been proposed, the trade-offs between safety and helpfulness, and their application to Large Vision-Language Models (LVLMs), are not w

Cited by 0SourcePDFScholar
2024

From LLMs to MLLMs: Exploring the Landscape of Multimodal Jailbreaking

EMNLP 2024main

The rapid development of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) has exposed vulnerabilities to various adversarial attacks. This paper provides a comprehensive overview of jailbreaking research targeting both LLMs and MLLMs, highlighting recent advancements in eval…

Cited by 3SourcePDFScholar