← Search

Wenhui Li

14 accepted papers

2026

Reason2Attack: Jailbreaking Text-to-Image Models via LLM Reasoning

AAAI 2026technical

Text-to-Image (T2I) models typically deploy safety mechanisms to prevent the generation of sensitive images. Unfortunately, recent jailbreaking attack methods manually design instructions for the LLM to generate adversarial prompts, which effectively exposing safety vulnerabilities of T2I models.

Cited by 0SourcePDFScholar
2026

SGoT-R1: Social Graph of Thought Reasoning-Enhanced Multimodal Large Language Model for Harmful Meme Detection

AAAI 2026technical

Internet memes serve as widely distributed multimodal social content that conveys complex ideas through metaphorical expressions, often containing harmful implications that make accurate harmful meme detection an important problem. Reasoning knowledge extracted from large language models plays a cru

Cited by 0SourcePDFScholar
2026

T2I-RiskyPrompt: A Benchmark for Safety Evaluation, Attack, and Defense on Text-to-Image Model

AAAI 2026technical

Using risky text prompts, such as pornography and violent prompts, to test the safety of text-to-image (T2I) models is a critical task. However, existing risky prompt datasets are limited in three key areas: 1) limited risky categories, 2) coarse-grained annotation, and 3) low effectiveness. To addr

Cited by 0SourcePDFScholar
2026

Thinking as Society: Multi-Social-Agent Self-Distillation for Multimodal Misinformation Detection

ICLR 2026poster

Multimodal Misinformation Detection (MMD) in realistic, mixed-sourced scenarios must incorporate robust reasoning capabilities to handle the social complexity and diverse types of forgeries. While MLLM-based agents are increasingly used for MMD task due to their powerful reasoning abilities, they su…

Cited by 0SourceScholar
2025

Global Static Pruning via Adaptive Sample Complexity Awareness

ICASSP 2025accepted

Dynamic pruning leverage the feature information of each input sample to dynamically adjust the network structure, generating multiple subnetworks suitable for different sample complexity. However, it inevitably introduces higher computational complexity and increased memory consumption. In addition…

Cited by 0SourceScholar
2025

JointSwinUNETR: an Efficient Feature-enhanced Architecture for Small Intestine Cine MRI Segmentation

ICASSP 2025accepted

The Cine MRI of the small intestine is a dynamic magnetic resonance imaging technique used to observe and evaluate small intestine motility. It captures sequential images of the organ in motion over time through rapid imaging. The Transformer architecture is highly effective at capturing long-range…

Cited by 0SourceScholar
2025

MncCap: Mining Neural Composition for Zero-shot Image Captioning via Text-only Training

ICASSP 2025accepted

Current text-only image captioning methods leverage the shared feature space of CLIP to train zero-shot image captioning using text data only, leaving feature associations and contextual understanding not fully explored. Neurological studies have revealed that the anterior temporal lobes of the brai…

Cited by 0SourceScholar
2024

MVITP: Multi-View Image-Text Perception for Few-Shot Remote Sensing Image Classification

ICASSP 2024accepted

Few-shot learning has been extensively applied in current remote sensing image classification, enabling rapid identification of new classes by leveraging prior knowledge effectively. However, current methods mainly rely on image modality to address the issue of low intra-class similarity and high in…

Cited by 0SourceScholar
2023

Mx2M: Masked Cross-Modality Modeling in Domain Adaptation for 3D Semantic Segmentation

AAAI 2023technical

Existing methods of cross-modal domain adaptation for 3D semantic segmentation predict results only via 2D-3D complementarity that is obtained by cross-modal feature matching. However, as lacking supervision in the target domain, the complementarity is not always reliable. The results are not ideal…

Cited by 6SourcePDFScholar
2022

Learning Memory-Augmented Unidirectional Metrics for Cross-Modality Person Re-Identification

CVPR 2022poster

This paper tackles the cross-modality person re-identification (re-ID) problem by suppressing the modality discrepancy. In cross-modality re-ID, the query and gallery images are in different modalities. Given a training identity, the popular deep classification baseline shares the same proxy (i.e.,…

Cited by 178PDFScholar
2022

Memory-Based Jitter: Improving Visual Recognition on Long-Tailed Data with Diversity in Memory

AAAI 2022technical

This paper considers deep visual recognition on long-tailed data. To make our method general, we tackle two applied scenarios, i.e. , deep classification and deep metric learning. Under the long-tailed data distribution, the most classes (i.e., tail classes) only occupy relatively few samples and ar…

Cited by 42SourcePDFScholar
2020

Consistent Domain Structure Learning and Domain Alignment for 2D Image-Based 3D Objects Retrieval

IJCAI 2020poster

2D image-based 3D objects retrieval is a new topic for 3D objects retrieval which can be used to manage 3D data with 2D images. The goal is to search some related 3D objects when given a 2D image. The task is challenging due to the large domain gap between 2D images and 3D objects. Therefore, it is…

Cited by 0SourcePDFScholar
2020

Deep Representation Learning on Long-Tailed Data: A Learnable Embedding Augmentation Perspective

CVPR 2020poster

This paper considers learning deep features from long-tailed data. We observe that in the deep feature space, the head classes and the tail classes present different distribution patterns. The head classes have a relatively large spatial span, while the tail classes have a significantly small spatia…

Cited by 287PDFScholar
2020

Hierarchical Instance Feature Alignment for 2D Image-Based 3D Shape Retrieval

IJCAI 2020poster

2D image-based 3D shape retrieval has become a hot research topic since its wide industrial applications and academic significance. However, existing view-based 3D shape retrieval methods are restricted by two settings, 1) learn the common-class features while neglecting the instance visual charac…

Cited by 0SourcePDFScholar