← Search

Jiyuan Jia

3 accepted papers

2025

VQAGuider: Guiding Multimodal Large Language Models to Answer Complex Video Questions

ACL 2025long

Complex video question-answering (VQA) requires in-depth understanding of video contents including object and action recognition as well as video classification and summarization, which exhibits great potential in emerging applications in education and entertainment, etc. Multimodal large language m…

2024

HOTVCOM: Generating Buzzworthy Comments for Videos

ACL 2024findings

In the era of social media video platforms, popular “hot-comments” play a crucial role in attracting user impressions of short-form videos, making them vital for marketing and branding purpose. However, existing research predominantly focuses on generating descriptive comments or “danmaku” in Englis…

2023

MaskFusion: Feature Augmentation for Click-Through Rate Prediction via Input-adaptive Mask Fusion

ICLR 2023poster

Click-through rate (CTR) prediction plays important role in the advertisement, recommendation, and retrieval applications. Given the feature set, how to fully utilize the information from the feature set is an active topic in deep CTR model designs. There are several existing deep CTR works focusing…

Cited by 2SourcePDFScholar