← Search

Weihong Du

5 accepted papers

2026

Unsupervised Semantic Discovery via Global and Local Semantic Alignment in Multimodal Clustering

AAAI 2026technical

Unsupervised multimodal semantic discovery aims to learn discriminative representations from multimodal data. However, existing methods suffer from two key limitations. First, they only align instances across modalities without modeling semantic-level consistency, which fails to mitigate semantic bi

Cited by 0SourcePDFScholar
2025

BAR: A Backward Reasoning based Agent for Complex Minecraft Tasks

ACL 2025finding

Large language model (LLM) based agents have shown great potential in following human instructions and automatically completing various tasks. To complete a task, the agent needs to decompose it into easily executed steps by planning. Existing studies mainly conduct the planning by inferring what st…

2024

CARE: A Clue-guided Assistant for CSRs to Read User Manuals

ACL 2024long

It is time-saving to build a reading assistant for customer service representations (CSRs) when reading user manuals, especially information-rich ones. Current solutions don’t fit the online custom service scenarios well due to the lack of attention to user questions and possible responses. Hence, w…

2024

PAGED: A Benchmark for Procedural Graphs Extraction from Documents

ACL 2024long

Automatic extraction of procedural graphs from documents creates a low-cost way for users to easily understand a complex procedure by skimming visual graphs. Despite the progress in recent studies, it remains unanswered: whether the existing studies have well solved this task (Q1) and whether the em…

2023

Knowing-how & Knowing-that: A New Task for Machine Comprehension of User Manuals

ACL 2023findings

The machine reading comprehension (MRC) of user manuals has huge potential in customer service. However, current methods have trouble answering complex questions. Therefore, we introduce the knowing-how & knowing-that task that requires the model to answer factoid-style, procedure-style, and inconsi…