← Search

Xinxing Zu

2 accepted papers

2025

Incorporating Dense Knowledge Alignment into Unified Multimodal Representation Models

CVPR 2025poster

Leveraging Large Language Models (LLMs) for text representation has achieved significant success, but the exploration of using Multimodal LLMs (MLLMs) for multimodal representation remains limited. Previous MLLM-based representation studies have primarily focused on unifying the embedding space whil…

Cited by 0SourcePDFScholar
2024

VK-G2T: Vision and Context Knowledge Enhanced Gloss2text

ICASSP 2024accepted

Existing sign language translation methods follow a two-stage pipeline: first converting the sign language video to a gloss sequence (i.e., Sign2Gloss) and then translating the generated gloss sequence into a spoken language sentence (i.e., Gloss2Text). While previous studies have focused on boostin…

Cited by 0SourceScholar