← Search

Wenjia Xu

5 accepted papers

2026

Dissecting Multimodal In-Context Learning: Modality Asymmetries and Circuit Dynamics in modern Transformers

ICML 2026spotlight

Transformer-based multimodal large language models often exhibit in-context learning (ICL) capabilities. Motivated by this phenomenon, we ask: how do transformers learn to associate information across modalities from in-context examples? We investigate this through controlled experiments on small tr…

Cited by 1SourceScholar
2022

Learning Prototype via Placeholder for Zero-shot Recognition

IJCAI 2022poster

Zero-shot learning (ZSL) aims to recognize unseen classes by exploiting semantic descriptions shared between seen classes and unseen classes. Current methods show that it is effective to learn visual-semantic alignment by projecting semantic embeddings into the visual space as class prototypes. Ho…

2022

VGSE: Visually-Grounded Semantic Embeddings for Zero-Shot Learning

CVPR 2022poster

Human-annotated attributes serve as powerful semantic embeddings in zero-shot learning. However, their annotation process is labor-intensive and needs expert supervision. Current unsupervised semantic embeddings, i.e., word embeddings, enable knowledge transfer between classes. However, word embeddi…

Cited by 76PDFcodeScholar
2020

Attribute Prototype Network for Zero-Shot Learning

NeurIPS 2020poster

From the beginning of zero-shot learning research, visual attributes have been shown to play an important role. In order to better transfer attribute-based knowledge from known to unknown classes, we argue that an image representation with integrated attribute localization ability would be beneficia…

Cited by 378SourcePDFScholar
2020

Compare and Reweight: Distinctive Image Captioning Using Similar Images Sets

ECCV 2020poster

A wide range of image captioning models has been developed, achieving significant improvement based on popular metrics, such as BLEU, CIDEr, and SPICE. However, although the generated captions can accurately describe the image, they are generic for similar images and lack distinctiveness, i.e., cann…

Cited by 52SourcePDFScholar