← Search

Jianfu Zhang

27 accepted papers

2026

AVGGT: Rethinking Global Attention for Accelerating VGGT

CVPR 2026

Models such as VGGT and \pi^3 have shown strong multi-view 3D performance, but their heavy reliance on global self-attention results in high computational cost. Existing sparse-attention variants offer partial speedups, yet lack a systematic analysis of how global attention contributes to multi-view

Cited by 0SourceScholar
2026

CVF-DLO: Cross-Visual-Field Branched Deformable Linear Objects Route Estimation

ICRA 2026poster

The perception of deformable linear objects (DLOs) poses significant challenges in robotic manipulation. Crossovers, mergings, and bifurcations of multiple DLOs complicate the identification of individual DLO physical instances. Furthermore, DLOs are often too large to be captured by a single camera…

Cited by 0codeScholar
2026

FakeXplain: AI-Generated Image Detection via Human-Aligned Grounded Reasoning

ICLR 2026poster

The rapid rise of image generation calls for detection methods that are both interpretable and reliable. Existing approaches, though accurate, act as black boxes and fail to generalize to out-of-distribution data, while multi-modal large language models (MLLMs) provide reasoning ability but often ha…

Cited by 0SourcecodeScholar
2026

High-Quality Full-Head 3D Avatar Generation from Any Single Portrait Image

AAAI 2026technical

In this work, we introduce a novel high-fidelity full-head 3D avatar generation method from a single image, regardless of perspective, style, expression, or accessories. Prior works often fail to preserve consistent head geometry and facial details, primarily due to their limited capacity in modelin

Cited by 0SourcePDFScholar
2026

Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images

CVPR 2026

The rapid growth of AI-generated imagery has blurred the boundary between real and synthetic content, raising practical concerns for digital integrity. Vision-language models (VLMs) can provide natural language explanations, but standard one-pass classifiers often miss subtle artifacts in high-quali

Cited by 0SourceScholar
2026

VTONGuard: Automatic Detection and Authentication of AI-Generated Virtual Try-On Content

ICASSP 2026poster

With the rapid advancement of generative AI, virtual try-on (VTON) systems are becoming increasingly common in e-commerce and digital entertainment. However, the growing realism of AI-generated try-on content raises pressing concerns about authenticity and responsible use. To address this, we presen…

Cited by 0SourcePDFScholar
2025

CVF-DLO: Cross-Visual-Field Branched Deformable Linear Objects Route Estimation

RA-L 2025

The perception of deformable linear objects (DLOs) poses significant challenges in robotic manipulation. Crossovers, mergings, and bifurcations of multiple DLOs complicate the identification of individual DLO physical instances. Furthermore, DLOs are often too large to be captured by a single camera

Cited by 1SourcecodeScholar
2025

GeneMAN: Generalizable Single-Image 3D Human Reconstruction from Multi-Source Human Data

NeurIPS 2025poster

Given a single in-the-wild human photo, it remains a challenging task to reconstruct a high-fidelity 3D human model. Existing methods face difficulties including a) the varying body proportions captured by in-the-wild human images; b) diverse personal belongings within the shot; and c) ambiguities i…

Cited by 0SourceScholar
2025

Pedestrian Motion Reconstruction: A Large-scale Benchmark via Mixed Reality Rendering with Multiple Perspectives and Modalities

ICLR 2025poster

Reconstructing pedestrian motion from dynamic sensors, with a focus on pedestrian intention, is crucial for advancing autonomous driving safety. However, this task is challenging due to data limitations arising from technical complexities, safety, and cost concerns. We introduce the Pedestrian Motio…

Cited by 0SourcePDFScholar
2025

WildFake: A Large-Scale and Hierarchical Dataset for AI-Generated Images Detection

AAAI 2025technical

The development of text-to-image generative models has enabled the creation of images so realistic that distinguishing between AI-generated images and real photos is becoming a challenge. This progress offers new possibilities but also raises concerns over privacy, authenticity, and security. Detect…

2024

DomainGallery: Few-shot Domain-driven Image Generation by Attribute-centric Finetuning

NeurIPS 2024poster

The recent progress in text-to-image models pretrained on large-scale datasets has enabled us to generate various images as long as we provide a text prompt describing what we want. Nevertheless, the availability of these models is still limited when we expect to generate images that fall into a spe…

2024

Hierarchical Attacks on Large-Scale Graph Neural Networks

ICASSP 2024accepted

In this paper, we present a novel hierarchical approach to adversarial attacks targeting Graph Neural Networks (GNNs), tailored to overcome the complexities inherent in large-scale poisoning attacks. Traditional global attack strategies often fail to yield effective results on extensive graph struct…

Cited by 0SourceScholar
2024

ProAug: Prototype-Based Augmentation for Long-Tailed Image Classification

ICASSP 2024accepted

Real-world data often exhibit long-tailed distributions with heavy class imbalance, which deteriorates the generalization performance of the classifier. To mitigate this problem, we propose a novel Prototype-based Augmentation framework (ProAug) to address the data scarcity issue by augmenting the f…

Cited by 0SourceScholar
2023

Amodal Instance Segmentation via Prior-Guided Expansion

AAAI 2023technical

Amodal instance segmentation aims to infer the amodal mask, including both the visible part and occluded part of each object instance. Predicting the occluded parts is challenging. Existing methods often produce incomplete amodal boxes and amodal masks, probably due to lacking visual evidences to ex…

Cited by 9SourcePDFScholar
2023

Critical Firms Prediction for Stemming Contagion Risk in Networked-Loans through Graph-Based Deep Reinforcement Learning

AAAI 2023technical

The networked-loan is major financing support for Micro, Small and Medium-sized Enterprises (MSMEs) in some developing countries. But external shocks may weaken the financial networks' robustness; an accidental default may spread across the network and collapse the whole network. Thus, predicting th…

Cited by 4SourcePDFScholar
2022

DeltaGAN: Towards Diverse Few-Shot Image Generation with Sample-Specific Delta

ECCV 2022poster

"Learning to generate new images for a novel category based on only a few images, named as few-shot image generation, has attracted increasing research interest. Several state-of-the-art works have yielded impressive results, but the diversity is still limited. In this work, we propose a novel Delta…

2022

SAFA: Sample-Adaptive Feature Augmentation for Long-Tailed Image Classification

ECCV 2022poster

"Imbalanced datasets with long-tailed distribution widely exist in practice, posing great challenges for deep networks on how to handle the biased predictions between head (majority, frequent) classes and tail (minority, rare) classes. Feature space of tail classes learned by deep networks is usuall…

Cited by 29SourcePDFScholar
2021

Parallel Multi-Resolution Fusion Network for Image Inpainting

ICCV 2021poster

Conventional deep image inpainting methods are based on auto-encoder architecture, in which the spatial details of images will be lost in the down-sampling process, leading to the degradation of generated results. Also, the structure information in deep layers and texture information in shallow laye…

Cited by 39PDFScholar
2020

DoveNet: Deep Image Harmonization via Domain Verification

CVPR 2020poster

Image composition is an important operation in image processing, but the inconsistency between foreground and background significantly degrades the quality of composite image. Image harmonization, aiming to make the foreground compatible with the background, is a promising yet challenging task. Howe…

Cited by 260PDFcodeScholar