← Search

Miaoge Li

11 accepted papers

2026

STiTch: Semantic Transition and Transportation in Collaboration for Training-Free Zero-Shot Composed Image Retrieval

CVPR 2026

Training-free zero-shot composed image retrieval models are recently gaining increasing research interest due to their generalizability and flexibility in unseen multimodal retrieval. Recent LLM-based advances focus on generating the expected target caption by exploring the compositional ability beh

Cited by 0SourceScholar
2025

Dynamic Multimodal Prototype Learning in Vision-Language Models

ICCV 2025poster

With the increasing attention to pre-trained vision-language models (VLMs), e.g., CLIP, substantial efforts have been devoted to many downstream tasks, especially in test-time adaptation (TTA). However, previous works focus on learning prototypes only in the textual modality while overlooking the am…

Cited by 0SourcePDFScholar
2025

Exploring Transferable Homogenous Groups for Compositional Zero-Shot Learning

IJCAI 2025

Conditional dependency present one of the trickiest problems in Compositional Zero-Shot Learning, leading to significant property variations of the same state (object) across different objects (states). To address this problem, existing approaches often adopt either all-to-one or one-to-one represen

2025

TsCA: On the Semantic Consistency Alignment via Conditional Transport for Compositional Zero-Shot Learning

IJCAI 2025

Compositional Zero-Shot Learning (CZSL) aims to recognize novel state-object compositions by leveraging the shared knowledge of their primitive components. Despite considerable progress, effectively calibrating the bias between semantically similar multimodal representations, as well as generalizing

2024

Instruction Tuning-free Visual Token Complement for Multimodal LLMs

ECCV 2024poster

"As the open community of large language models (LLMs) matures, multimodal LLMs (MLLMs) have promised an elegant bridge between vision and language. However, current research is inherently constrained by challenges such as the need for high-quality instruction pairs and the loss of visual informatio…

Cited by 3SourcePDFScholar
2024

Patch-Prompt Aligned Bayesian Prompt Tuning for Vision-Language Models

UAI 2024poster

For downstream applications of vision-language pre-trained models, there has been significant interest in constructing effective prompts. Existing works on prompt engineering, which either require laborious manual designs or optimize the prompt tuning as a point estimation problem, may fail to descr…

Cited by 3SourcePDFScholar
2023

PatchCT: Aligning Patch Set and Label Set with Conditional Transport for Multi-Label Image Classification

ICCV 2023poster

Multi-label image classification is a prediction task that aims to identify more than one label from a given image. This paper considers the semantic consistency of the latent space between the visual patch and linguistic label domains and introduces the conditional transport (CT) theory to bridge t…

Cited by 22PDFcodeScholar
2023

Tuning Multi-mode Token-level Prompt Alignment across Modalities

NeurIPS 2023poster

Advancements in prompt tuning of vision-language models have underscored their potential in enhancing open-world visual concept comprehension. However, prior works only primarily focus on single-mode (only one prompt for each modality) and holistic level (image or sentence) semantic alignment, which…

2022

Knowledge-Aware Bayesian Deep Topic Model

NeurIPS 2022accept

We propose a Bayesian generative model for incorporating prior domain knowledge into hierarchical topic modeling. Although embedded topic models (ETMs) and its variants have gained promising performance in text analysis, they mainly focus on mining word co-occurrence patterns, ignoring potentially e…