← Search

Daqing Liu

10 accepted papers

2025

Scaling Down Text Encoders of Text-to-Image Diffusion Models

CVPR 2025poster

Text encoders in diffusion models have rapidly evolved, transitioning from CLIP to T5-XXL. Although this evolution has significantly enhanced the models' ability to understand complex prompts and generate text, it also leads to a substantial increase in the number of parameters. Despite T5 series en…

2024

Decomposing Semantic Shifts for Composed Image Retrieval

AAAI 2024technical

Composed image retrieval is a type of image retrieval task where the user provides a reference image as a starting point and specifies a text on how to shift from the starting point to the desired target image. However, most existing methods focus on the composition learning of text and reference im…

2023

Cocktail: Mixing Multi-Modality Control for Text-Conditional Image Generation

NeurIPS 2023poster

Text-conditional diffusion models are able to generate high-fidelity images with diverse contents. However, linguistic representations frequently exhibit ambiguous descriptions of the envisioned objective imagery, requiring the incorporation of additional control signals to bolster the efficacy of t…

Cited by 23SourcePDFScholar
2023

Exploring Temporal Concurrency for Video-Language Representation Learning

ICCV 2023poster

Paired video and language data is naturally temporal concurrency, which requires the modeling of the temporal dynamics within each modality and the temporal alignment across modalities simultaneously. However, most existing video-language representation learning methods only focus on discrete semant…

Cited by 4PDFcodeScholar
2023

Modeling Video As Stochastic Processes for Fine-Grained Video Representation Learning

CVPR 2023highlight

A meaningful video is semantically coherent and changes smoothly. However, most existing fine-grained video representation learning methods learn frame-wise features by aligning frames across videos or exploring relevance between multiple views, neglecting the inherent dynamic process of each video.…

2022

Modeling Image Composition for Complex Scene Generation

CVPR 2022poster

We present a method that achieves state-of-the-art results on challenging (few-shot) layout-to-image generation tasks by accurately modeling textures, structures and relationships contained in a complex scene. After compressing RGB images into patch tokens, we propose the Transformer with Focal Atte…

Cited by 57PDFcodeScholar
2022

SemMAE: Semantic-Guided Masking for Learning Masked Autoencoders

NeurIPS 2022accept

Recently, significant progress has been made in masked image modeling to catch up to masked language modeling. However, unlike words in NLP, the lack of semantic decomposition of images still makes masked autoencoding (MAE) different between vision and language. In this paper, we explore a potential…

2020

Learning to Discretely Compose Reasoning Module Networks for Video Captioning

IJCAI 2020poster

Generating natural language descriptions for videos, i.e., video captioning, essentially requires step-by-step reasoning along the generation process. For example, to generate the sentence “a man is shooting a basketball”, we need to first locate and describe the subject “man”, next reason out the m…

2020

More Grounded Image Captioning by Distilling Image-Text Matching Model

CVPR 2020poster

Visual attention not only improves the performance of image captioners, but also serves as a visual interpretation to qualitatively measure the caption rationality and model transparency. Specifically, we expect that a captioner can fix its attentive gaze on the correct objects while generating the…

Cited by 179PDFcodeScholar