← Search

Hewen Pan

3 accepted papers

2026

UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Models

CVPR 2026

With the advancement of multi-modal Large Language Models (LLMs), Video LLMs have been further developed to perform on holistic and specialized video understanding. However, existing works are limited to specialized video understanding tasks, failing to achieve a comprehensive and multi-grained vide

Cited by 0SourceScholar
2025

AdvEDM: Fine-grained Adversarial Attack against VLM-based Embodied Agents

NeurIPS 2025poster

Vision-Language Models (VLMs), with their strong reasoning and planning capabilities, are widely used in embodied decision-making (EDM) tasks in embodied agents, such as autonomous driving and robotic manipulation. Recent research has increasingly explored adversarial attacks on VLMs to reveal their…

Cited by 0SourceScholar
2025

An Efficient Residual-based Low-dose PET Reconstruction with Spatial-Frequency Integration

ICASSP 2025accepted

Positron emission tomography (PET) is a nuclear medical imaging technique where image quality depends on the dose of radionuclides administered to the patient. While standard-dose PET (SPET) offers high-quality imaging, it also poses radiation risks. If reconstructing low-dose PET (LPET) images can…

Cited by 0SourceScholar