← Search

Haoming Zhou

3 accepted papers

2026

Cross Modal Fine-grained Alignment via Granularity-aware and Region-uncertain Modeling

AAAI 2026technical

Fine-grained image-text alignment is a pivotal challenge in multimodal learning, underpinning key applications such as visual question answering, image captioning, and vision-language navigation. Unlike global alignment, fine-grained alignment requires precise correspondence between localized visual

Cited by 0SourcePDFScholar
2024

Decoupled Self-Adaptive Distribution Regularization for Few-Shot Image Classification

ICASSP 2024accepted

The feature dispersion, arising from the inherent constraints of data scarcity, has emerged as a prominent challenge in the domain of few-shot learning. In this paper, we propose a novel Self-adaptive Distribution Regularization (SADR) approach, which can adaptively bridge the semantic gaps across d…

Cited by 0SourceScholar
2021

Kaleido-BERT: Vision-Language Pre-Training on Fashion Domain

CVPR 2021poster

We present a new vision-language (VL) pre-training model dubbed Kaleido-BERT, which introduces a novel kaleido strategy for fashion cross-modality representations from transformers. In contrast to random masking strategy of recent VL models, we design alignment guided masking to jointly focus more o…

Cited by 154PDFcodeScholar