← Search

Jianping Fan

28 accepted papers

2026

Leveraging Evidence Priors for Robust Prompt Learning under Noisy Supervision in Vision-Language Models

ICML 2026poster

Prompt learning for vision-language models (VLMs) often suffers from performance degradation when adapting to downstream tasks with noisy labels. Existing methods that rely on filtering or reconstructing supervision can propagate errors, leading to sharp performance drops. We observe that pre-traine…

Cited by 0SourceScholar
2025

Analyzing the Effects of Supervised Fine-Tuning on Model Knowledge from Token and Parameter Levels

EMNLP 2025

Large language models (LLMs) acquire substantial world knowledge during pre-training, which is further shaped by post-training techniques such as supervised fine-tuning (SFT). However, the impact of SFT on a model’s knowledge remains underexplored, limiting our ability to control knowledge behavior

Cited by 0SourcePDFScholar
2025

Attention Mechanisms Perspective: Exploring LLM Processing of Graph-Structured Data

ICML 2025poster

Attention mechanisms are critical to the success of large language models (LLMs), driving significant advancements in multiple fields. However, for graph-structured data, which requires emphasis on topological connections, they fall short compared to message-passing mechanisms on fixed links, such a…

2025

DoGA: Enhancing Grounded Object Detection via Grouped Pre-Training with Attributes

AAAI 2025technical

Recent advances in vision-language pre-training have significantly enhanced the model capabilities on grounded object detection. However, these studies often pre-train with coarse-grained text prompts, such as plain category names and brief grounded phrases. This limitation curtails the model's capa…

2025

Instruct-of-Reflection: Enhancing Large Language Models Iterative Reflection Capabilities via Dynamic-Meta Instruction

NAACL 2025long

Self-reflection for Large LanguageModels (LLMs) has gained significant attention. Existing approaches involve models iterating and improving their previous responses based on LLMs’ internal reflection ability or external feedback. However, recent research has raised doubts about whether intrinsic se…

2025

Iterative Self-Training with Class-Aware Text-to-Image Synthesis for Visual Task Learning

AAAI 2025technical

Generative models are widely used to produce synthetic images with annotations, alleviating the burden of image collection and annotation for training deep visual models. However, challenges such as limited image diversity, noisy pseudo labels, and domain gaps between synthetic and real images often…

Cited by 0SourcePDFScholar
2025

Less Attention is More: Prompt Transformer for Generalized Category Discovery

CVPR 2025poster

Generalized Category Discovery (GCD) typically relies on the pre-trained Vision Transformer (ViT) to extract features from a global receptive field, followed by contrastive learning to simultaneously classify unlabeled known classes and unknown classes without priors. Owing to the deficiency in the…

2025

MAPS: Motivation-Aware Personalized Search via LLM-Driven Consultation Alignment

ACL 2025long

Personalized product search aims to retrieve and rank items that match users’ preferences and search intent. Despite their effectiveness, existing approaches typically assume that users’ query fully captures their real motivation. However, our analysis of a real-world e-commerce platform reveals tha…

2025

Omni-Query Active Learning for Source-Free Domain Adaptive Cross-Modality 3D Semantic Segmentation

AAAI 2025technical

Source-Free Domain Adaptation (SFDA) aims to transfer a pre-trained source model to the unlabeled target domain without accessing the source data, thereby effectively solving labeled data dependency and domain shift problems. However, the SFDA setting faces a bottleneck due to the absence of supervi…

2025

Open-Unfairness Adversarial Mitigation for Generalized Deepfake Detection

ICCV 2025poster

Deepfake detection methods are becoming increasingly crucial for identity security and have recently been employed to support legal proceedings. However, these methods often exhibit unfairness due to flawed logical reasoning, undermining the reliability of their predictions and raising concerns abou…

2025

Similarity = Value? Consultation Value-Assessment and Alignment for Personalized Search

EMNLP 2025

Personalized search systems in e-commerce platforms increasingly involve user interactions with AI assistants, where users consult about products, usage scenarios, and more. Leveraging consultation to personalize search services is trending. Existing methods typically rely on semantic similarity to

2025

TL-Training: A Task-Feature-Based Framework for Training Large Language Models in Tool Use

EMNLP 2025

Large language models (LLMs) achieve remarkable advancements by leveraging tools to interact with environments, a critical step toward generalized AI. However, the standard supervised fine-tuning (SFT) approach, which relies on large-scale datasets, often overlooks task-specific characteristics in t

2024

Image Clustering with External Guidance

ICML 2024oral

The core of clustering lies in incorporating prior knowledge to construct supervision signals. From classic k-means based on data compactness to recent contrastive clustering guided by self-supervision, the evolution of clustering methods intrinsically corresponds to the progression of supervision s…

2023

ANetQA: A Large-Scale Benchmark for Fine-Grained Compositional Reasoning Over Untrimmed Videos

CVPR 2023poster

Building benchmarks to systemically analyze different capabilities of video question answering (VideoQA) models is challenging yet crucial. Existing benchmarks often use non-compositional simple questions and suffer from language biases, making it difficult to diagnose model weaknesses incisively. A…

2023

Dual Pseudo-Labels Interactive Self-Training for Semi-Supervised Visible-Infrared Person Re-Identification

ICCV 2023poster

Visible-infrared person re-identification (VI-ReID) aims to match a specific person from a gallery of images captured from non-overlapping visible and infrared cameras. Most works focus on fully supervised VI-ReID, which requires substantial cross-modality annotation that is more expensive than the…

Cited by 42PDFcodeScholar
2023

Learning How to Learn Domain-Invariant Parameters for Domain Generalization

ICASSP 2023accepted

Due to domain shift, deep neural networks (DNNs) usually fail to generalize well on unknown test data in practice. Domain generalization (DG) aims to overcome this issue by capturing domain-invariant representations from source domains. Motivated by the insight that only partial parameters of DNNs a…

Cited by 0SourceScholar
2023

Long-Tailed Recognition with Causal Invariant Transformation

ICASSP 2023accepted

Standard classification models rely on the assumption that all the classes of interest are equally represented in training datasets. However, visual phenomena exhibit a long-tailed distribution, such that many standard approaches fail to properly model and result in a considerable degeneration on ac…

Cited by 0SourceScholar
2023

SAP-DETR: Bridging the Gap Between Salient Points and Queries-Based Transformer Detector for Fast Model Convergency

CVPR 2023poster

Recently, the dominant DETR-based approaches apply central-concept spatial prior to accelerating Transformer detector convergency. These methods gradually refine the reference points to the center of target objects and imbue object queries with the updated central reference information for spatially…

2023

VS-Boost: Boosting Visual-Semantic Association for Generalized Zero-Shot Learning

IJCAI 2023poster

Unlike conventional zero-shot learning (CZSL) which only focuses on the recognition of unseen classes by using the classifier trained on seen classes and semantic embeddings, generalized zero-shot learning (GZSL) aims at recognizing both the seen and unseen classes, so it is more challenging due to…

Cited by 17SourcePDFScholar
2022

Class Guided Channel Weighting Network for Fine-Grained Semantic Segmentation

AAAI 2022technical

Deep learning has achieved promising performance on semantic segmentation, but few works focus on semantic segmentation at the fine-grained level. Fine-grained semantic segmentation requires recognizing and distinguishing hundreds of sub-categories. Due to the high similarity of different sub-catego…

Cited by 2SourcePDFScholar
2022

Cross-Domain Few-Shot Learning for Rare-Disease Skin Lesion Segmentation

ICASSP 2022accepted

Recently, deep learning (DL)-based skin lesion segmentation in dermoscopic images has advanced the efficient diagnosis of skin diseases. Commonly, most of the DL-based methods require a large amount of training data and can only perform accurate predictions on pre-defined classes. However, there exi…

Cited by 0SourceScholar
2022

U-GAT-VC: Unsupervised Generative Attentional Networks for Non-Parallel Voice Conversion

ICASSP 2022accepted

Non-parallel voice conversion (VC) is a technique of transfer-ring voice from one style to another without using a parallel corpus in model training. Various methods are proposed to approach non-parallel VC using deep neural networks. Among them, CycleGAN-VC and its variants have been widely accepte…

Cited by 0SourceScholar
2021

DeepME: Deep Mixture Experts for Large-scale Image Classification

IJCAI 2021poster

Although deep learning has demonstrated its outstanding performance on image classification, most well-known deep networks make efforts to optimize both their structures and their node weights for recognizing fewer (e.g., no more than 1000) object classes. Therefore, it is attractive to extend or mi…

Cited by 4SourcePDFScholar
2020

Adaptive Fractional Dilated Convolution Network for Image Aesthetics Assessment

CVPR 2020poster

To leverage deep learning for image aesthetics assessment, one critical but unsolved issue is how to seamlessly incorporate the information of image aspect ratios to learn more robust models. In this paper, an adaptive fractional dilated convolution (AFDC), which is aspect-ratio-embedded, compositio…

Cited by 113PDFScholar
2020

Learning Deep Network for Detecting 3D Object Keypoints and 6D Poses

CVPR 2020poster

The state-of-art 6D object pose detection methods use convolutional neural networks to estimate objects' 6D poses from RGB images. However, they require huge numbers of images with explicit 3D annotations such as 6D poses, 3D bounding boxes and 3D keypoints, either obtained by manual labeling or inf…

Cited by 38PDFScholar
2017

Multi-Modal Factorized Bilinear Pooling With Co-Attention Learning for Visual Question Answering

ICCV 2017poster

Visual question answering (VQA) is challenging because it requires a simultaneous understanding of both the visual content of images and the textual content of questions. The approaches used to represent the images and questions in a fine-grained manner and questions and to fuse these multi-modal fe…

Cited by 885PDFcodeScholar