← Search

Dong Jing

5 accepted papers

2026

Say Cheese! Detail-Preserving Portrait Collection Generation via Natural Language Edits

CVPR 2026

As social media platforms proliferate, users increasingly demand intuitive ways to create diverse, high-quality portrait collections. In this work, we introduce Portrait Collection Generation (PCG), a novel task that generates coherent portrait collections by editing a reference portrait image throu

Cited by 0SourceScholar
2025

Leveraging Large Vision-Language Model as User Intent-Aware Encoder for Composed Image Retrieval

AAAI 2025technical

Composed Image Retrieval (CIR) aims to retrieve target images from candidate set using a hybrid-modality query consisting of a reference image and a relative caption that describes the user intent. Recent studies attempt to utilize Vision-Language Pre-training Models (VLPMs) with various fusion stra…

Cited by 2SourcePDFScholar
2024

FineCLIP: Self-distilled Region-based CLIP for Better Fine-grained Understanding

NeurIPS 2024poster

Contrastive Language-Image Pre-training (CLIP) achieves impressive performance on tasks like image classification and image-text retrieval by learning on large-scale image-text datasets. However, CLIP struggles with dense prediction tasks due to the poor grasp of the fine-grained details. Although e…

Cited by 3SourcePDFScholar
2021

Saliency-based Multi-View Mixed Language Training for Zero-shot Cross-lingual Classification

EMNLP 2021finding

Recent multilingual pre-trained models, like XLM-RoBERTa (XLM-R), have been demonstrated effective in many cross-lingual tasks. However, there are still gaps between the contextualized representations of similar words in different languages. To solve this problem, we propose a novel framework named…

Cited by 8SourcePDFScholar