← Search

Bingchan Zhao

3 accepted papers

2025

Exploring Fine-Grained Human Motion Video Captioning

COLING 2025main

Detailed descriptions of human motion are crucial for effective fitness training, which highlights the importance of research in fine-grained human motion video captioning. Existing video captioning models often fail to capture the nuanced semantics of videos, resulting in the generated descriptions…

2025

ISR: Self-Refining Referring Expressions for Entity Grounding

ACL 2025long

Entity grounding, a crucial task in constructing multimodal knowledge graphs, aims to align entities from knowledge graphs with their corresponding images. Unlike conventional visual grounding tasks that use referring expressions (REs) as inputs, entity grounding relies solely on entity names and ty…

2025

MPO: Boosting LLM Agents with Meta Plan Optimization

EMNLP 2025

Recent advancements in large language models (LLMs) have enabled LLM-based agents to successfully tackle interactive planning tasks. However, despite their successes, existing approaches often suffer from planning hallucinations and require retraining for each new agent. To address these challenges,