← Search

Zhiqiang Yuan

12 accepted papers

2026

F2RVLM: Boosting Fine-grained Fragment Retrieval for Multi-Modal Long-form Dialogue with Vision Language Model

AAAI 2026technical

Traditional dialogue retrieval aims to select the most appropriate utterance or image from recent dialogue history. However, they often fail to meet users’ actual needs for revisiting semantically coherent content scattered across long-form conversations. To fill this gap, we define the Fine-grained

Cited by 0SourcePDFScholar
2026

Less Redundancy: Boosting Practicality of Vision Language Model in Walking Assistants

ICASSP 2026oral

Approximately 283 million people worldwide live with visual impairments, motivating increasing research into leveraging Visual Language Models (VLMs) to develop effective walking assistance systems for blind and low vision individuals. However, existing VLMs in walking assistant task often have outp…

Cited by 0SourcePDFScholar
2025

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets

ICCV 2025poster

Vision-language models (VLMs) often struggle with compositional reasoning due to insufficient high-quality image-text data. To tackle this challenge, we propose a novel block-based diffusion approach that automatically generates counterfactual datasets without manual annotation. Our method utilizes…

Cited by 0SourcePDFScholar
2025

From Imitation to Innovation: The Emergence of AI's Unique Artistic Styles and the Challenge of Copyright Protection

ICCV 2025poster

Current legal frameworks consider AI-generated works eligible for copyright protection when they meet originality requirements and involve substantial human intellectual input. However, systematic legal standards and reliable evaluation methods for AI art copyrights are lacking. Through comprehensiv…

Cited by 0SourcePDFScholar
2025

ILDiff: Generate Transparent Animated Stickers by Implicit Layout Distillation

ICASSP 2025accepted

High-quality animated stickers usually contain transparent channels, which are often ignored by current video generation models. To generate fine-grained animated transparency channels, existing methods can be roughly divided into video matting algorithms and diffusion-based algorithms. The methods…

Cited by 0SourceScholar
2025

MCID: Multi-aspect Copyright Infringement Detection for Generated Images

ICCV 2025poster

With the rapid advancement of generative models, we can now create highly realistic images. This represents a significant technical breakthrough but also introduces new challenges for copyright protection. Previous methods for detecting copyright infringement in AI-generated images mainly depend on…

Cited by 0SourcePDFScholar
2025

Secret Lies in Color: Enhancing AI-Generated Images Detection with Color Distribution Analysis

CVPR 2025poster

The advancement of Generative Adversarial Networks (GANs) and diffusion models significantly enhances the realism of synthetic images, driving progress in image processing and creative design. However, this progress also necessitates the development of effective detection methods, as synthetic image…

Cited by 0SourcePDFScholar
2025

Semantic to Structure: Learning Structural Representations for Infringement Detection

ICASSP 2025accepted

Structural information in images is crucial for aesthetic assessment, and it is widely recognized in the artistic field that imitating the structure of other works significantly infringes on creators’ rights. The advancement of diffusion models has led to AI-generated content imitating artists’ stru…

Cited by 0SourceScholar
2025

WalkVLM: Aid Visually Impaired People Walking by Vision Language Model

ICCV 2025poster

Approximately 200 million individuals around the world suffer from varying degrees of visual impairment, making it crucial to leverage AI technology to offer walking assistance for these people.With the recent progress of vision-language models (VLMs), applying VLMs to offer walking guidance has bec…

Cited by 0SourcePDFScholar
2021

Pixel Contrastive-Consistent Semi-Supervised Semantic Segmentation

ICCV 2021poster

We present a novel semi-supervised semantic segmentation method which jointly achieves two desiderata of segmentation model regularities: the label-space consistency property between image augmentations and the feature-space contrastive property among different pixels. We leverage the pixel-level L2…

Cited by 226PDFScholar
2021

Reaching Pruning Locations in a Vine Using a Deep Reinforcement Learning Policy

ICRA 2021poster

We outline a neural network-based pipeline for perception, control and planning of a 7 DoF robot for tasks that involve reaching into a dormant grapevine canopy. The proposed system consists of a 6 DoF industrial robot arm and a linear slider that can actuate on an entire grape vine. Our approach us…

Cited by 15SourceScholar
2021

Semantically Robust Unpaired Image Translation for Data With Unmatched Semantics Statistics

ICCV 2021poster

Many applications of unpaired image-to-image translation require the input contents to be preserved semantically during translations. Unaware of the inherently unmatched semantics distributions between source and target domains, existing distribution matching methods (i.e., GAN-based) can give undes…

Cited by 28PDFcodeScholar