← Search

Xuguang Duan

4 accepted papers

2024

DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation

ICLR 2024poster

Subject-driven text-to-image generation aims to generate customized images of the given subject based on the text descriptions, which has drawn increasing attention. Existing methods mainly resort to finetuning a pretrained generative model, where the identity-relevant information (e.g., the boy) an…

2023

Decouple Before Interact: Multi-Modal Prompt Learning for Continual Visual Question Answering

ICCV 2023poster

In the real world, a desirable Visual Question Answering model is expected to provide correct answers to new questions and images in a continual setting (recognized as CL-VQA). However, existing works formulate CLVQA from a vision-only or language-only perspective, and straightforwardly apply the un…

Cited by 24PDFScholar
2022

Parametric Visual Program Induction with Function Modularization

ICML 2022spotlight

Generating programs to describe visual observations has gained much research attention recently. However, most of the existing approaches are based on non-parametric primitive functions, making them unable to handle complex visual scenes involving many attributes and details. In this paper, we propo…

Cited by 3SourcePDFScholar
2018

Weakly Supervised Dense Event Captioning in Videos

NeurIPS 2018poster

Dense event captioning aims to detect and describe all events of interest contained in a video. Despite the advanced development in this area, existing methods tackle this task by making use of dense temporal annotations, which is dramatically source-consuming. This paper formulates a new problem: w…