← Search

Yanxin Long

2 accepted papers

2025

DialogGen: Multi-modal Interactive Dialogue System with Multi-turn Text-Image Generation

NAACL 2025findings

Text-to-image (T2I) generation models have significantly advanced in recent years. However, effective interaction with these models is challenging for average users due to the need for specialized prompt engineering knowledge and the inability to perform multi-turn image generation, hindering a dyna…

2023

NLIP: Noise-Robust Language-Image Pre-training

AAAI 2023technical

Large-scale cross-modal pre-training paradigms have recently shown ubiquitous success on a wide range of downstream tasks, e.g., zero-shot classification, retrieval and image captioning. However, their successes highly rely on the scale and quality of web-crawled data that naturally contain much inc…

Cited by 33SourcePDFScholar