← Search

Fei Su

21 accepted papers

2026

Adapting In-context Generation for Enhanced Composed Image Retrieval

CVPR 2026

As a challenge vision-language task, Composed Image Retrieval (CIR) aims to integrate information from a bi-modal query (image + text) to retrieve target images. While supervised CIR has achieved notable success in domain-specific scenarios, its reliance on manually annotated triplets restricts its

Cited by 0SourcecodeScholar
2026

Inter-Edit: First Benchmark for Interactive Instruction-Based Image Editing

CVPR 2026

Precise and controllable image editing remains a significant challenge. Current methods often rely on text prompts, but achieving accurate spatial localization solely through descriptions is inherently difficult. Mask-based approaches, though offering better control, typically require overly precise

Cited by 0SourcecodeScholar
2026

Modality and Task Adaptation for Enhanced Zero-shot Composed Image Retrieval

AAAI 2026technical

As a challenging vision-language task, Zero-Shot Composed Image Retrieval (ZS-CIR) is designed to retrieve target images using bi-modal (image+text) queries. Typical ZS-CIR methods employ an inversion network to generate pseudo-word tokens that effectively represent the input semantics. However, the

Cited by 0SourcePDFScholar
2026

RAA: Achieving Interactive Remove/Add Anything via Fully Synthetic Data

AAAI 2026technical

Precise and controllable image editing, especially object removal and insertion, represents one of the most common demands in image manipulation. However, existing methods suffer from severe limitations. Mask-based inpainting often introduces visual artifacts and semantic inconsistencies, while inst

Cited by 0SourcePDFScholar
2026

Rethinking Two-Stage Referring-by-Tracking in Referring Multi-Object Tracking: Make it Strong Again

CVPR 2026

Referring Multi-Object Tracking (RMOT) aims to track multiple objects specified by natural language expressions in videos. With the recent significant progress of one-stage methods, the two-stage Referring-by-Tracking (RBT) paradigm has gradually lost its popularity. However, its lower training cost

Cited by 0SourcecodeScholar
2025

Automatic Synthetic Data and Fine-grained Adaptive Feature Alignment for Composed Person Retrieval

NeurIPS 2025poster

Person retrieval has attracted rising attention. Existing methods are mainly divided into two retrieval modes, namely image-only and text-only. However, they are unable to make full use of the available information and are difficult to meet diverse application requirements. To address the above limi…

Cited by 0SourcecodeScholar
2025

Filter or Compensate: Towards Invariant Representation from Distribution Shift for Anomaly Detection

AAAI 2025technical

Recent Anomaly Detection (AD) methods have achieved great success with In-Distribution (ID) data. However, real-world data often exhibits distribution shift, causing huge performance decay on traditional AD methods. From this perspective, few previous work has explored AD with distribution shift, an…

2025

R1-ShareVL: Incentivizing Reasoning Capabilities of Multimodal Large Language Models via Share-GRPO

NeurIPS 2025poster

In this work, we aim to incentivize the reasoning ability of Multimodal Large Language Models (MLLMs) via reinforcement learning (RL) and develop an effective approach that mitigates the sparse reward and advantage vanishing issues during RL. To this end, we propose Share-GRPO, a novel RL approach…

Cited by 0SourcecodeScholar
2025

UFO: Enhancing Diffusion-Based Video Generation with a Uniform Frame Organizer

AAAI 2025technical

Recently, diffusion-based video generation models have achieved significant success. However, existing models often suffer from issues like weak consistency and declining image quality over time. To overcome these challenges, inspired by aesthetic principles, we propose a non-invasive plug-in called…

2024

Dynamic Clustering and Cluster Contrastive Learning for Unsupervised Person Re-Id With Feature Distribution Alignment

ICASSP 2024accepted

Unsupervised Re-ID methods aim at learning robust and discriminative features from unlabeled data. However, existing methods often ignore the noise from distribution discrepancy during network training, which may lead to feature misalignment and hinder the model performance. To address this problem,…

Cited by 0SourceScholar
2024

PracticalDG: Perturbation Distillation on Vision-Language Models for Hybrid Domain Generalization

CVPR 2024poster

Domain Generalization (DG) aims to resolve distribution shifts between source and target domains and current DG methods are default to the setting that data from source and target domains share identical categories. Nevertheless there exists unseen classes from target domains in practical scenarios.…

2021

Diagnosing Covid-19 from CT Images Based on an Ensemble Learning Framework

ICASSP 2021accepted

Research on automated diagnosis of Coronavirus Disease 2019 (COVID-19) has increased in recent months. SPGC COVID19 aims at classifying the grouped images of the same patient into COVID, Community Acquired Pneumonia(CAP) or normal. In this paper, we propose a novel ensemble learning framework to sol…

Cited by 0SourceScholar
2021

SANet++: Enhanced Scale Aggregation with Densely Connected Feature Fusion for Crowd Counting

ICASSP 2021accepted

Crowd counting has gained considerable attention recently but remains challenging mainly due to large scale variations. In this paper, we present SANet++ with a novel architecture to generate high-quality density maps and further perform accurate counting. SANet++ obtains enhanced multi-scale repres…

Cited by 0SourceScholar
2020

Adaptive Elastic Loss Based on Progressive Inter-Class Association for Cervical Histology Image Segmentation

ICASSP 2020accepted

Cervical cancer is one of the most commonly diagnosed cancer types worldwide, while is curable if detected early. However, few computer-aided algorithms have been explored on cervical histology image, which is vital for abnormality assessment. In this paper, an end-to-end deep segmentation network f…

Cited by 0SourceScholar
2019

Jointly Predicting Future Sequence and Steering Angles for Dynamic Driving Scenes

ICASSP 2019accepted

Generative Adversarial Network (GAN) has attracted rising attention for video future sequence prediction in driving scenes. However, the images generated by GAN often miss the target for lack of any constraints for its generated target. In this paper, an encoder-decoder based multi-task video predic…

Cited by 0SourceScholar
2018

Scale Aggregation Network for Accurate and Efficient Crowd Counting

ECCV 2018poster

In this paper, we propose a novel encoder-decoder network, called extit{Scale Aggregation Network (SANet)}, for accurate and efficient crowd counting. The encoder extracts multi-scale features with scale aggregation modules and the decoder generates high-resolution density maps by using a set of tra…

Cited by 835SourcePDFScholar