← Search

Min Song

4 accepted papers

2026

TIPO: Text to Image with Text Pre-sampling for Prompt Optimization

ICLR 2026poster

TIPO (Text-to-Image Prompt Optimization) introduces an efficient approach for automatic prompt refinement in text-to-image (T2I) generation. Starting from simple user prompts, TIPO leverages a lightweight pre-trained model to expand these prompts into richer, detailed versions. Conceptually, TIPO sa…

Cited by 0SourceScholar
2025

Def-DTS: Deductive Reasoning for Open-domain Dialogue Topic Segmentation

ACL 2025finding

Dialogue Topic Segmentation (DTS) aims to divide dialogues into coherent segments. DTS plays a crucial role in various NLP downstream tasks, but suffers from chronic problems: data shortage, labeling ambiguity, and incremental complexity of recently proposed solutions. On the other hand, Despite adv…

2025

Retrieval Visual Contrastive Decoding to Mitigate Object Hallucinations in Large Vision-Language Models

ACL 2025finding

Despite significant advancements in Large Vision-Language Models, Object Hallucination (OH) remains a persistent challenge. Building upon prior studies on contrastive decoding that address this issue without requiring additional model training, we introduce RVCD (Retrieval Visual Contrastive Decodin…

2024

TelME: Teacher-leading Multimodal Fusion Network for Emotion Recognition in Conversation

NAACL 2024long

Emotion Recognition in Conversation (ERC) plays a crucial role in enabling dialogue sys- tems to effectively respond to user requests. The emotions in a conversation can be identi- fied by the representations from various modal- ities, such as audio, visual, and text. How- ever, due to the weak cont…