← Search

Benjamin Yao

3 accepted papers

2024

Open Vocabulary Multi-Label Video Classification

ECCV 2024poster

"Pre-trained vision-language models (VLMs) have enabled significant progress in open vocabulary computer vision tasks such as image classification, object detection and image segmentation. Some recent works have focused on extending VLMs to open vocabulary single label action classification in video…

Cited by 2SourcePDFScholar
2024

X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs

ECCV 2024poster

"Recent advancements in Multimodal Large Language Models (MLLMs) have revolutionized the field of vision-language understanding by integrating visual perception capabilities into Large Language Models (LLMs). The prevailing trend in this field involves the utilization of a vision encoder derived fro…

Cited by 2SourcePDFScholar
2022

Joint Goal Segmentation and Goal Success Prediction on Multi-Domain Conversations

COLING 2022main

To evaluate the performance of a multi-domain goal-oriented Dialogue System (DS), it is important to understand what the users’ goals are for the conversations and whether those goals are successfully achieved. The success rate of goals directly correlates with user satisfaction and perceived useful…

Cited by 2SourcePDFScholar