← Search

Zhongzhou Zhao

4 accepted papers

2025

Incorporating Dense Knowledge Alignment into Unified Multimodal Representation Models

CVPR 2025poster

Leveraging Large Language Models (LLMs) for text representation has achieved significant success, but the exploration of using Multimodal LLMs (MLLMs) for multimodal representation remains limited. Previous MLLM-based representation studies have primarily focused on unifying the embedding space whil…

Cited by 0SourcePDFScholar
2024

VK-G2T: Vision and Context Knowledge Enhanced Gloss2text

ICASSP 2024accepted

Existing sign language translation methods follow a two-stage pipeline: first converting the sign language video to a gloss sequence (i.e., Sign2Gloss) and then translating the generated gloss sequence into a spoken language sentence (i.e., Gloss2Text). While previous studies have focused on boostin…

Cited by 0SourceScholar
2023

COOP: Decoupling and Coupling of Whole-Body Grasping Pose Generation

ICCV 2023poster

Generating life-like whole-body human grasping has garnered significant attention in the field of computer graphics. Existing works have demonstrated the effectiveness of keyframe-guided motion generation framework, witch focus on modeling the grasping motions of humans in temporal sequence when the…

Cited by 7PDFcodeScholar
2023

Towards Zero-Shot Personalized Table-to-Text Generation with Contrastive Persona Distillation

ICASSP 2023accepted

Existing neural methods have shown great potentials towards generating informative text from structured tabular data as well as maintaining high content fidelity. However, few of them shed light on generating personalized expressions, which often requires well-aligned persona-table-text datasets tha…

Cited by 0SourceScholar