← Search

Ziwei Zheng

5 accepted papers

2026

DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning

ICLR 2026poster

Large Vision-Language Models excel at multimodal understanding but struggle to deeply integrate visual information into their predominantly text-based reasoning processes, a key challenge in mirroring human cognition. To address this, we introduce DeepEyes, a model that learns to ``think with images…

Cited by 0SourcecodeScholar
2026

HeadHunt-VAD: Hunting Robust Anomaly-Sensitive Heads in MLLM for Tuning-Free Video Anomaly Detection

AAAI 2026technical

Video Anomaly Detection (VAD) aims to locate events that deviate from normal patterns in videos. Traditional approaches often rely on extensive labeled data and incur high computational costs. Recent tuning-free methods based on Multimodal Large Language Models (MLLMs) offer a promising alternative

Cited by 0SourcePDFScholar
2025

Nullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace Projection

CVPR 2025poster

Recent studies have shown that large vision-language models (LVLMs) often suffer from the issue of object hallucinations (OH). To mitigate this issue, we introduce an efficient method that edits the model weights based on an unsafe subspace, which we call HalluSpace in this paper. With truthful and…

2024

DyFADet: Dynamic Feature Aggregation for Temporal Action Detection

ECCV 2024poster

"Recent proposed neural network-based Temporal Action Detection (TAD) models are inherently limited to extracting the discriminative representations and modeling action instances with various lengths from complex scenes by shared-weights detection heads. Inspired by the successes in dynamic neural n…

2024

Fine-grained Dynamic Network for Generic Event Boundary Detection

ECCV 2024poster

"Generic event boundary detection (GEBD) aims at pinpointing event boundaries naturally perceived by humans, playing a crucial role in understanding long-form videos. Given the diverse nature of generic boundaries, spanning different video appearances, objects, and actions, this task remains challen…