DMGIN: How Multimodal LLMs Enhance Large Recommendation Models for Lifelong User Post-click Behaviors
Zhuoxing Wei, Qingchen Xie, Qi Liu, Jingsong Yu
Abstract
Modeling user interest based on lifelong user behavior sequences is crucial for enhancing Click-Through Rate (CTR) prediction. However, long post-click behavior sequences themselves pose severe performance issues: the sheer volume of data leads to high computational costs and inefficiencies in model training and inference. Traditional methods address this by introducing two-stage approaches, but this compromises model effectiveness due to incomplete utilization of the full sequence context. More importantly, integrating multimodal embeddings into existing large recommendation models (LRM) presents significant challenges: These embeddings often exacerbate computational burdens and mismatch with LRM architectures. To address these issues and enhance the model
BibTeX
@inproceedings{aaai2026_dmginhowmultimod,
title = {DMGIN: How Multimodal LLMs Enhance Large Recommendation Models for Lifelong User Post-click Behaviors},
author = {Zhuoxing Wei and Qingchen Xie and Qi Liu and Jingsong Yu},
booktitle = {AAAI 2026},
year = {2026}
}