← Search

Yiwei Sun

4 accepted papers

2026

From Evaluation to Defense: Advancing Safety in Video Large Language Models

ICLR 2026poster

While the safety risks of image-based large language models (Image LLMs) have been extensively studied, their video-based counterparts (Video LLMs) remain critically under-examined. To systematically study this problem, we introduce \textbf{VideoSafetyEval} - the first large-scale, real-world benchm…

Cited by 0SourceScholar
2026

Meerkat-VL: Implicit Risk Safety Alignment in Multimodal LLMs via Perceptual Reasoning and Self-Verification

ICML 2026poster

Multimodal LLMs (MLLMs) are increasingly deployed across diverse applications, but they pose significant safety concerns due to cross-modal interactions. To improve model safety awareness, existing methods rely on explicit-risk preference datasets and reinforcement learning guided by safety rewards.…

Cited by 0SourceScholar
2026

RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document Understanding

AAAI 2026technical

Multi-modal Retrieval-Augmented Generation (RAG) has become a critical method for empowering LLMs by leveraging candidate visual documents. However, current methods consider the entire document as the basic retrieval unit, introducing substantial irrelevant visual content in two ways: 1) Relevant do

Cited by 0SourcePDFScholar