← Search

Fanyi Wang

8 accepted papers

2025

Overcoming Heterogeneous Data in Federated Medical Vision-Language Pre-training: A Triple-Embedding Model Selector Approach

AAAI 2025technical

The scarcity data of medical field brings the collaborative training in medical vision-language pre-training (VLP) cross different clients. Therefore, the collaborative training in medical VLP faces two challenges: First, the medical data requires privacy, thus can not directly shared across differe…

2024

BARET: Balanced Attention Based Real Image Editing Driven by Target-Text Inversion

AAAI 2024technical

Image editing approaches with diffusion models have been rapidly developed, yet their applicability are subject to requirements such as specific editing types (e.g., foreground or background object editing, style transfer), multiple conditions (e.g., mask, sketch, caption), and time consuming fine-t…

Cited by 5SourcePDFScholar
2024

DBS: Differentiable Budget-Aware Searching For Channel Pruning

ICASSP 2024accepted

Network pruning is an effective technique to reduce computation costs for deep model deployment on resource-constraint devices. Searching superior sub-networks from a vast search space through Neural Architecture Search (NAS) , which conducts a one-shot supernet used as a performance estimator, is s…

Cited by 0SourceScholar
2024

IC-FPS: Instance-Centroid Faster Point Sampling Framework for 3D Point-based Object Detection

IROS 2024

3D object detection is one of the most important tasks in autonomous driving and robotics. Our research focuses on tackling low efficiency issue of point-based methods, and we propose a novel Instance-Centroid Faster Point Sampling (IC-FPS) framework. We design a Neighboring Feature Diffusion Module

Cited by 1SourceScholar
2024

Lightweight High-Resolution Subject Matting in the Real World

ICASSP 2024accepted

Existing saliency object detection (SOD) methods struggle to satisfy fast inference and accurate results simultaneously in high resolution scenes. They are limited by the quality of public datasets and efficient network modules for high-resolution images. To alleviate these issues, we propose to con…

Cited by 0SourceScholar
2024

Zero-shot High-fidelity and Pose-controllable Character Animation

IJCAI 2024poster

Image-to-video (I2V) generation aims to create a video sequence from a single image, which requires high temporal coherence and visual fidelity. However, existing approaches suffer from inconsistency of character appearances and poor preservation of fine details. Moreover, they require a large amoun…

Cited by 4SourcePDFScholar
2023

GAM: Gradient Attention Module of Optimization for Point Clouds Analysis

AAAI 2023technical

In the point cloud analysis task, the existing local feature aggregation descriptors (LFAD) do not fully utilize the neighborhood information of center points. Previous methods only use the distance information to constrain the local aggregation process, which is easy to be affected by abnormal poin…

2023

Matting Moments: A Unified Data-Driven Matting Engine for Mobile AIGC in Photo Gallery

IJCAI 2023poster

Image matting is a fundamental technique in visual understanding and has become one of the most significant capabilities in mobile phones. Despite the development of mobile storage and computing power, achieving diverse mobile Artificial Intelligence Generated Content (AIGC) applications remains a g…

Cited by 3SourcePDFScholar