← Search

Zhenfei Zhang

4 accepted papers

2026

FAVE: A Structured Benchmark for Fine-Grained Audio-Visual Temporal Evaluation in Multimodal LLMs

CVPR 2026

Audio-visual large language models (AVLLMs) have made significant strides in understanding visual and auditory content. However, their ability to capture fine-grained temporal relationships between audio and visual streams remains insufficiently evaluated. To address this, we introduce FAVE (Fine-gr

Cited by 0SourceScholar
2024

A New Benchmark and Model for Challenging Image Manipulation Detection

AAAI 2024technical

The ability to detect manipulation in multimedia data is vital in digital forensics. Existing Image Manipulation Detection (IMD) methods are mainly based on detecting anomalous features arisen from image editing or double compression artifacts. All existing IMD techniques encounter challenges when i…

2022

Improving Class Activation Map for Weakly Supervised Object Localization

ICASSP 2022accepted

We propose a Weakly Supervised Object Localization (WSOL) method that can locate an object within a given image using a pre-trained network learned with only class labels without location annotations. Most existing WSOL methods rely on thresholding a Class Activation Map (CAM) generated by the pre-t…

Cited by 0SourceScholar