← Search

Yajie Zhang

3 accepted papers

2025

LamRA: Large Multimodal Model as Your Advanced Retrieval Assistant

CVPR 2025poster

With the rapid advancement of multimodal information retrieval, increasingly complex retrieval tasks have emerged. Existing methods predominately rely on task-specific fine-tuning of vision-language models, often those trained with image-text contrastive learning. In this paper, we explore the possi…

Cited by 8SourcePDFScholar
2025

SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs

EMNLP 2025

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities by integrating visual and textual inputs, yet modality alignment remains one of the most challenging aspects. Current MLLMs typically rely on simple adapter architectures and pretraining approaches to bridge vision en

2019

Learning Latent Representations for Style Control and Transfer in End-to-end Speech Synthesis

ICASSP 2019accepted

In this paper, we introduce the Variational Autoencoder (VAE) to an end-to-end speech synthesis model, to learn the latent representation of speaking styles in an unsupervised manner. The style representation learned through VAE shows good properties such as disentangling, scaling, and combination,…

Cited by 0SourceScholar