← Search

Guoxing Yang

9 accepted papers

2025

Leveraging Large Vision-Language Model as User Intent-Aware Encoder for Composed Image Retrieval

AAAI 2025technical

Composed Image Retrieval (CIR) aims to retrieve target images from candidate set using a hybrid-modality query consisting of a reference image and a relative caption that describes the user intent. Recent studies attempt to utilize Vision-Language Pre-training Models (VLPMs) with various fusion stra…

Cited by 2SourcePDFScholar
2024

DIFFSC: Semantic Communication Framework With Enhanced Denoising Through Diffusion Probabilistic Models

ICASSP 2024accepted

In communication systems, the challenge of ensuring accurate data transmission across noisy channels remains paramount. While semantic communication shows potential in improving image transmission and reconstruction, existing methods still suffer from perceptual quality degradation in high-noise env…

Cited by 0SourceScholar
2024

FineCLIP: Self-distilled Region-based CLIP for Better Fine-grained Understanding

NeurIPS 2024poster

Contrastive Language-Image Pre-training (CLIP) achieves impressive performance on tasks like image classification and image-text retrieval by learning on large-scale image-text datasets. However, CLIP struggles with dense prediction tasks due to the poor grasp of the fine-grained details. Although e…

Cited by 3SourcePDFScholar
2024

Image Retrieval with Composed Query by Multi-Scale Multi-Modal Fusion

ICASSP 2024accepted

Image retrieval with composed query (IR-CQ) is a challenging task since it aims to retrieve the target image according to a hybrid-modality query which consists of a reference image and a text modifier. Previous approaches mainly focus on designing various multi-modal fusion modules to fuse the hybr…

Cited by 0SourceScholar
2024

Progressive Image Synthesis from Semantics to Details with Denoising Diffusion GAN

ICASSP 2024accepted

Although denoising diffusion probabilistic models (DDPMs) have shown remarkable progress in image generation, they typically face two main challenges: the time-expensive sampling process and the semantically meaningless latent space, which are often addressed separately in previous works. In particu…

Cited by 0SourceScholar
2024

UniAdapter: Unified Parameter-Efficient Transfer Learning for Cross-modal Modeling

ICLR 2024poster

Large-scale vision-language pre-trained models have shown promising transferability to various downstream tasks. As the size of these foundation models and the number of downstream tasks grow, the standard full fine-tuning paradigm becomes unsustainable due to heavy computational and storage costs.…

2024

VDT: General-purpose Video Diffusion Transformers via Mask Modeling

ICLR 2024poster

This work introduces Video Diffusion Transformer (VDT), which pioneers the use of transformers in diffusion-based video generation. It features transformer blocks with modularized temporal and spatial attention modules to leverage the rich spatial-temporal representation inherited in transformers. A…

2022

Visual Prompt Tuning for Few-Shot Text Classification

COLING 2022main

Deploying large-scale pre-trained models in the prompt-tuning paradigm has demonstrated promising performance in few-shot learning. Particularly, vision-language pre-training models (VL-PTMs) have been intensively explored in various few-shot downstream tasks. However, most existing works only apply…

2021

L2M-GAN: Learning To Manipulate Latent Space Semantics for Facial Attribute Editing

CVPR 2021poster

A deep facial attribute editing model strives to meet two requirements: (1) attribute correctness -- the target attribute should correctly appear on the edited face image; (2) irrelevance preservation -- any irrelevant information (e.g., identity) should not be changed after editing. Meeting both re…

Cited by 86PDFcodeScholar