← Search

Xiaoyong Zhu

9 accepted papers

2026

Cross-modal Identity Mapping: Minimizing Information Loss in Modality Conversion via Reinforcement Learning

CVPR 2026

Large Vision-Language Models (LVLMs) often omit or misrepresent critical visual content in generated image captions. Minimizing such information loss will force LVLMs to focus on image details to generate precise descriptions. However, measuring information loss during modality conversion is inheren

Cited by 0SourceScholar
2026

ORION: Decoupling and Alignment for Unified Autoregressive Understanding and Generation

ICLR 2026poster

Unified multimodal Large Language Models (MLLMs) hold great promise for seamlessly integrating understanding and generation. However, monolithic autoregressive architectures, despite their elegance and conversational fluency, suffer from a fundamental semantic–structural conflict: optimizing for low…

Cited by 0SourceScholar
2026

iTryOn: Mastering Interactive Video Virtual Try-On with Spatial-Semantic Guidance

ICML 2026poster

Video Virtual Try-On (VVT) aims to seamlessly replace a garment on a person in a video with a new one. While existing methods have made significant strides in maintaining temporal consistency, they are predominantly confined to non-interactive scenarios where models merely showcase garments. This li…

Cited by 0SourceScholar
2025

Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models

ACL 2025long

With the rapid advancement of Large Language Models (LLMs), significant safety concerns have emerged. Fundamentally, the safety of large language models is closely linked to the accuracy, comprehensiveness, and clarity of their understanding of safety knowledge, particularly in domains such as law,…

Cited by 0SourcePDFScholar
2025

HiddenDetect: Detecting Jailbreak Attacks against Multimodal Large Language Models via Monitoring Hidden States

ACL 2025long

The integration of additional modalities increases the susceptibility of large vision-language models (LVLMs) to safety risks, such as jailbreak attacks, compared to their language-only counterparts. While existing research primarily focuses on post-hoc alignment techniques, the underlying safety me…

Cited by 0SourcePDFScholar
2025

INTER: Mitigating Hallucination in Large Vision-Language Models by Interaction Guidance Sampling

ICCV 2025poster

Hallucinations in large vision-language models (LVLMs) pose significant challenges for real-world applications, as LVLMs may generate responses that appear plausible yet remain inconsistent with the associated visual content. This issue rarely occurs in human cognition. We argue that this discrepanc…

2025

PGPO: Enhancing Agent Reasoning via Pseudocode-style Planning Guided Preference Optimization

ACL 2025finding

Large Language Model (LLM) agents have demonstrated impressive capabilities in handling complex interactive problems. Existing LLM agents mainly generate natural language plans to guide reasoning, which is verbose and inefficient. NL plans are also tailored to specific tasks and restrict agents’ abi…

2025

See the World, Discover Knowledge: A Chinese Factuality Evaluation for Large Vision Language Models

ACL 2025finding

The evaluation of factual accuracy in large vision language models (LVLMs) has lagged behind their rapid development, making it challenging to fully reflect these models’ knowledge capacity and reliability. In this paper, we introduce the first factuality-based visual question-answering benchmark in…

Cited by 0SourcePDFScholar
2025

Towards Visualization-of-Thought Jailbreak Attack against Large Visual Language Models

NeurIPS 2025poster

As Visual Language Models (VLMs) continue to evolve, they have demonstrated increasingly sophisticated logical reasoning capabilities and multimodal thought generation, opening doors to widespread applications. However, this advancement raises serious concerns about content security, particularly wh…

Cited by 0SourcecodeScholar