← Search

Wenhua Zhang

7 accepted papers

2025

AIGuard: A Benchmark and Lightweight Detection for E-commerce AIGC Risks

ACL 2025finding

Recent advancements in AI-generated content (AIGC) have heightened concerns about harmful outputs, such as misinformation and malicious misuse.Existing detection methods face two key limitations:(1) lacking real-world AIGC scenarios and corresponding risk datasets, and(2) both traditional and multim…

2025

Adaptive Skeleton Prompt Tuning for Cross-Dataset 3D Human Pose Estimation

ICASSP 2025accepted

Inconsistency of distributions in human actions and camera viewpoints can lead to significant deviations when the pre-trained 3D pose estimators are tested on cross-datasets. In practical applications, the estimators usually follow the standard full fine-tuning paradigm on the target dataset, which…

Cited by 0SourceScholar
2025

Incorporating Dense Knowledge Alignment into Unified Multimodal Representation Models

CVPR 2025poster

Leveraging Large Language Models (LLMs) for text representation has achieved significant success, but the exploration of using Multimodal LLMs (MLLMs) for multimodal representation remains limited. Previous MLLM-based representation studies have primarily focused on unifying the embedding space whil…

Cited by 0SourcePDFScholar
2025

Multi-scale Feature Interaction and Adaptive Experts for Panoptic Segmentation in Remote Sensing Images

ICASSP 2025accepted

Panoptic segmentation unifies the traditional tasks of instance and semantic segmentation. It plays a crucial role in the field of remote sensing; however, it encounters challenges in recognizing small objects and in the model’s ability to generalize across complex scenes. In this paper, we introduc…

Cited by 0SourceScholar
2025

RFEM: Remote Feature Enhancement Module for Target Detection

ICASSP 2025accepted

The research and development of dense crowd detection technology have always been one of the hot and challenging topics in the field of computer vision. DETR-like models have shown good performance in both training efficiency and inference capabilities. Nevertheless, as the optimization proceeds, th…

Cited by 0SourceScholar
2024

Hybrid-SORT: Weak Cues Matter for Online Multi-Object Tracking

AAAI 2024technical

Multi-Object Tracking (MOT) aims to detect and associate all desired objects across frames. Most methods accomplish the task by explicitly or implicitly leveraging strong cues (i.e., spatial and appearance information), which exhibit powerful instance-level discrimination. However, when object occlu…

2024

MMM-RS: A Multi-modal, Multi-GSD, Multi-scene Remote Sensing Dataset and Benchmark for Text-to-Image Generation

NeurIPS 2024poster

Recently, the diffusion-based generative paradigm has achieved impressive general image generation capabilities with text prompts due to its accurate distribution modeling and stable training process. However, generating diverse remote sensing (RS) images that are tremendously different from general…