← Search

Delong Liu

7 accepted papers

2026

Adapting In-context Generation for Enhanced Composed Image Retrieval

CVPR 2026

As a challenge vision-language task, Composed Image Retrieval (CIR) aims to integrate information from a bi-modal query (image + text) to retrieve target images. While supervised CIR has achieved notable success in domain-specific scenarios, its reliance on manually annotated triplets restricts its

Cited by 0SourcecodeScholar
2026

Inter-Edit: First Benchmark for Interactive Instruction-Based Image Editing

CVPR 2026

Precise and controllable image editing remains a significant challenge. Current methods often rely on text prompts, but achieving accurate spatial localization solely through descriptions is inherently difficult. Mask-based approaches, though offering better control, typically require overly precise

Cited by 0SourcecodeScholar
2026

Modality and Task Adaptation for Enhanced Zero-shot Composed Image Retrieval

AAAI 2026technical

As a challenging vision-language task, Zero-Shot Composed Image Retrieval (ZS-CIR) is designed to retrieve target images using bi-modal (image+text) queries. Typical ZS-CIR methods employ an inversion network to generate pseudo-word tokens that effectively represent the input semantics. However, the

Cited by 0SourcePDFScholar
2026

RAA: Achieving Interactive Remove/Add Anything via Fully Synthetic Data

AAAI 2026technical

Precise and controllable image editing, especially object removal and insertion, represents one of the most common demands in image manipulation. However, existing methods suffer from severe limitations. Mask-based inpainting often introduces visual artifacts and semantic inconsistencies, while inst

Cited by 0SourcePDFScholar
2025

Automatic Synthetic Data and Fine-grained Adaptive Feature Alignment for Composed Person Retrieval

NeurIPS 2025poster

Person retrieval has attracted rising attention. Existing methods are mainly divided into two retrieval modes, namely image-only and text-only. However, they are unable to make full use of the available information and are difficult to meet diverse application requirements. To address the above limi…

Cited by 0SourcecodeScholar
2025

UFO: Enhancing Diffusion-Based Video Generation with a Uniform Frame Organizer

AAAI 2025technical

Recently, diffusion-based video generation models have achieved significant success. However, existing models often suffer from issues like weak consistency and declining image quality over time. To overcome these challenges, inspired by aesthetic principles, we propose a non-invasive plug-in called…