← Search

Quan Cui

9 accepted papers

2025

Oasis: One Image is All You Need for Multimodal Instruction Data Synthesis

ICCV 2025poster

The success of multi-modal large language models (MLLMs) has been largely attributed to the large-scale training data. However, the training data of many MLLMs is unavailable due to privacy concerns. The expensive and labor-intensive process of collecting multi-modal data further exacerbates the pro…

2024

Modeling Label Correlations with Latent Context for Multi-Label Recognition

ECCV 2024poster

"Label dependencies have been widely studied in multi-label image recognition for improving performances. Previous methods mainly considered label co-occurrences as label correlations. In this paper, we show that label co-occurrences may be insufficient to represent label correlations, and modeling…

Cited by 0SourcePDFScholar
2023

A Simple Framework for Text-Supervised Semantic Segmentation

CVPR 2023poster

Text-supervised semantic segmentation is a novel research topic that allows semantic segments to emerge with image-text contrasting. However, pioneering methods could be subject to specifically designed network architectures. This paper shows that a vanilla contrastive language-image pre-training (C…

2022

Contrastive Vision-Language Pre-training with Limited Resources

ECCV 2022poster

"Pioneering dual-encoder pre-training works (e.g., CLIP and ALIGN) have revealed the potential of aligning multi-modal representations with contrastive learning. However, these works require a tremendous amount of data and computational resources (e.g., billion-level web data and hundreds of GPUs),…

2022

Discriminability-Transferability Trade-Off: An Information-Theoretic Perspective

ECCV 2022poster

"This work simultaneously considers the discriminability and transferability properties of deep representations in the typical supervised learning task, i.e., image classification. By a comprehensive temporal analysis, we observe a trade-off between these two properties. The discriminability keeps i…

2020

BBN: Bilateral-Branch Network With Cumulative Learning for Long-Tailed Visual Recognition

CVPR 2020oral

Our work focuses on tackling the challenging but natural visual recognition task of long-tailed data distribution (i.e., a few classes occupy most of the data, while most classes have rarely few samples). In the literature, class re-balancing strategies (e.g., re-weighting and re-sampling) are the p…

Cited by 1043PDFcodeScholar
2020

ExchNet: A Unified Hashing Network for Large-Scale Fine-Grained Image Retrieval

ECCV 2020poster

Retrieving content relevant images from a large-scale fine-grained dataset could suffer from intolerably slow query speed and highly redundant storage cost, due to high-dimensional real-valued embeddings which aim to distinguish subtle visual differences of fine-grained objects. In this paper, we st…

Cited by 52SourcePDFScholar