← Search

Xinfeng Liao

2 accepted papers

2026

ModalSyncSum: Synchronizing Image and Text for Reliable Summary Generation

AAAI 2026technical

Multimodal summarization with multimodal output (MSMO) aims to generate coherent textual summaries while selecting the most semantically relevant images to enhance expressiveness. Despite the advancements of large multimodal models like GPT-4o, LLaMA-3, and Grok-3, these models often exhibit halluci

Cited by 0SourcePDFScholar
2024

An Effective Deployment of Diffusion LM for Data Augmentation in Low-Resource Sentiment Classification

EMNLP 2024main

Sentiment classification (SC) often suffers from low-resource challenges such as domain-specific contexts, imbalanced label distributions, and few-shot scenarios. The potential of the diffusion language model (LM) for textual data augmentation (DA) remains unexplored, moreover, textual DA methods st…