← Search

Zhen Xie

6 accepted papers

2026

Anatomical Prior-Driven Framework for Autonomous Robotic Cardiac Ultrasound Standard View Acquisition

ICRA 2026poster

Cardiac ultrasound diagnosis is critical for cardiovascular disease assessment, but acquiring standard views remains highly operator-dependent. Existing medical segmentation models often yield anatomically inconsistent results in images with poor textural differentiation between distinct feature cla…

2025

AirCache: Activating Inter-modal Relevancy KV Cache Compression for Efficient Large Vision-Language Model Inference

ICCV 2025poster

Recent advancements in Large Visual Language Models (LVLMs) have gained significant attention due to their remarkable reasoning capabilities and proficiency in generalization. However, processing a large number of visual tokens and generating long-context outputs impose substantial computational ove…

Cited by 0SourcePDFScholar
2025

ChartM3: A Multi-Stage Code-Driven Pipeline for Constructing Multi-Dimensional and Multi-Step Visual Reasoning Data in Chart Comprehension

EMNLP 2025

Complex chart understanding tasks demand advanced visual recognition and reasoning capabilities from multimodal large language models (MLLMs). However, current research provides limited coverage of complex chart scenarios and computation-intensive reasoning tasks prevalent in real-world applications

Cited by 0SourcePDFScholar
2024

IVTP: Instruction-guided Visual Token Pruning for Large Vision-Language Models

ECCV 2024poster

"Inspired by the remarkable achievements of Large Language Models (LLMs), Large Vision-Language Models (LVLMs) have likewise experienced significant advancements. However, the increased computational cost and token budget occupancy associated with lengthy visual tokens pose significant challenge to…

Cited by 3SourcePDFScholar
2020

SSTNet: Detecting Manipulated Faces Through Spatial, Steganalysis and Temporal Features

ICASSP 2020accepted

Compared to conventional object detection which focuses on high-level image content, face manipulation detection pays more attention to low-level artifacts and temporal discrepancies. However, there are few methods considering both of these two characteristics. In this work, we propose a novel manip…

Cited by 0SourceScholar