← Search

Shu Guo

5 accepted papers

2026

RMAdapter: Reconstruction-based Multi-Modal Adapter for Vision-Language Models

AAAI 2026technical

Pre-trained Vision-Language Models (VLMs), e.g. CLIP, have become essential tools in multimodal transfer learning. However, fine-tuning VLMs in few-shot scenarios poses significant challenges in balancing task-specific adaptation and generalization in the obtained model. Meanwhile, current researc

Cited by 0SourcePDFScholar
2025

Towards S²-Challenges Underlying LLM-Based Augmentation for Personalized News Recommendation

AAAI 2025technical

Personalized news recommendation aims to recommend candidate news to the target user. Since the data and knowledge involved in traditional recommender systems are restricted, recent studies utilize large language models (LLMs) to generate news articles and augment the original dataset. However, desp…

Cited by 0SourcePDFScholar
2025

Variational Multi-Modal Hypergraph Attention Network for Multi-Modal Relation Extraction

IJCAI 2025

Multi-modal relation extraction (MMRE) is a challenging task that seeks to identify relationships between entities with textual and visual attributes. However, existing methods struggle to handle the complexities posed by multiple entity pairs within a single sentence that share similar contextual i

2023

Dual-Gated Fusion with Prefix-Tuning for Multi-Modal Relation Extraction

ACL 2023findings

Multi-Modal Relation Extraction (MMRE) aims at identifying the relation between two entities in texts that contain visual clues. Rich visual content is valuable for the MMRE task, but existing works cannot well model finer associations among different modalities, failing to capture the truly helpful…

2023

Multi-Modal Knowledge Graph Transformer Framework for Multi-Modal Entity Alignment

EMNLP 2023long findings

Multi-Modal Entity Alignment (MMEA) is a critical task that aims to identify equivalent entity pairs across multi-modal knowledge graphs (MMKGs). However, this task faces challenges due to the presence of different types of information, including neighboring entities, multi-modal attributes, and ent…

Cited by 0SourcecodeScholar