← Search

Zhijiang Li

3 accepted papers

2026

Safe Autoregressive Image Generation with Iterative Self-Improving Codebooks

ICML 2026poster

Unlike diffusion-based models that operate in continuous latent spaces, autoregressive unified multimodal models produce images by sequentially predicting discretized visual tokens. These tokens are derived from a codebook that maps embeddings to quantized visual patterns. The language-like architec…

Cited by 0SourceScholar
2025

LLM Jailbreak Detection for (Almost) Free!

EMNLP 2025

Large language models (LLMs) enhance security through alignment when widely used, but remain susceptible to jailbreak attacks capable of producing inappropriate content. Jailbreak detection methods show promise in mitigating jailbreak attacks through the assistance of other models or multiple model

2025

Reimagining Safety Alignment with An Image

EMNLP 2025

Large language models (LLMs) excel in diverse applications but face dual challenges: generating harmful content under jailbreak attacks and over-refusing benign queries due to rigid safety mechanisms. These issues severely affect the application of LLMs, especially in the medical and education field