← Search

Jun Wan

20 accepted papers

2026

Cross-Slice Knowledge Transfer via Masked Multi-Modal Heterogeneous Graph Contrastive Learning for Spatial Gene Expression Inference

CVPR 2026

While spatial transcriptomics (ST) has advanced our understanding of gene expression in tissue context, its high experimental cost limits its large-scale application. Predicting ST from pathology images is a promising, cost-effective alternative, but existing methods struggle to capture complex cros

Cited by 0SourcecodeScholar
2026

SpaCRD: Multimodal Deep Fusion of Histology and Spatial Transcriptomics for Cancer Region Detection

AAAI 2026technical

Accurate detection of cancer tissue regions (CTR) enables deeper analysis of the tumor microenvironment and offers crucial insights into treatment response. Traditional CTR detection methods, which typically rely on the rich cellular morphology in histology images, are susceptible to a high rate of

Cited by 0SourcePDFScholar
2026

Veritas: Generalizable Deepfake Detection via Pattern-Aware Reasoning

ICLR 2026oral

Deepfake detection remains a formidable challenge due to the evolving nature of fake content in real-world scenarios. However, existing benchmarks suffer from severe discrepancies from industrial practice, typically featuring homogeneous training sources and low-quality testing images, which hinder…

Cited by 0SourcecodeScholar
2026

VideoVeritas: AI-Generated Video Detection via Perception Pretext Reinforcement Learning

ICML 2026poster

The growing capability of video generation poses escalating security risks, making reliable detection increasingly essential. In this paper, we introduce **VideoVeritas**, a framework that integrates fine-grained perception and fact-based reasoning. We observe that while current multi-modal large la…

Cited by 0SourceScholar
2025

Mixture-of-Attack-Experts with Class Regularization for Unified Physical-Digital Face Attack Detection

AAAI 2025technical

Unified detection of digital and physical attacks in facial recognition systems has become a focal point of research in recent years. However, current multi-modal methods typically ignore the intra-class and inter-class variability across different types of attacks, leading to degraded performance.…

Cited by 0SourcePDFScholar
2025

Recover and Match: Open-Vocabulary Multi-Label Recognition through Knowledge-Constrained Optimal Transport

CVPR 2025poster

Identifying multiple novel classes in an image, known as open-vocabulary multi-label recognition, is a challenging task in computer vision. Recent studies explore the transfer of powerful vision-language models such as CLIP. However, these approaches face two critical challenges: (1) The local seman…

2024

BP4ER: Bootstrap Prompting for Explicit Reasoning in Medical Dialogue Generation

COLING 2024main

Medical dialogue generation (MDG) has gained increasing attention due to its substantial practical value. Previous works typically employ a sequence-to-sequence framework to generate medical responses by modeling dialogue context as sequential text with annotated medical entities. While these method…

2024

CFPL-FAS: Class Free Prompt Learning for Generalizable Face Anti-spoofing

CVPR 2024highlight

Domain generalization (DG) based Face Anti-Spoofing (FAS) aims to improve the model's performance on unseen domains. Existing methods either rely on domain labels to align domain-invariant feature spaces or disentangle generalizable features from the whole sample which inevitably lead to the distort…

Cited by 34SourcePDFScholar
2024

Compound Text-Guided Prompt Tuning via Image-Adaptive Cues

AAAI 2024technical

Vision-Language Models (VLMs) such as CLIP have demonstrated remarkable generalization capabilities to downstream tasks. However, existing prompt tuning based frameworks need to parallelize learnable textual inputs for all categories, suffering from massive GPU memory consumption when there is a lar…

2024

Factorized Learning Assisted with Large Language Model for Gloss-free Sign Language Translation

COLING 2024main

Previous Sign Language Translation (SLT) methods achieve superior performance by relying on gloss annotations. However, labeling high-quality glosses is a labor-intensive task, which limits the further development of SLT. Although some approaches work towards gloss-free SLT through jointly training…

Cited by 14SourcePDFScholar
2024

Unified Physical-Digital Face Attack Detection

IJCAI 2024poster

Face Recognition (FR) systems can suffer from physical (i.e., print photo) and digital (i.e., DeepFake) attacks. However, previous related work rarely considers both situations at the same time. This implies the deployment of multiple models and thus more computational burden. The main reasons for t…

Cited by 15SourcePDFScholar
2024

VL-FAS: Domain Generalization via Vision-Language Model For Face Anti-Spoofing

ICASSP 2024accepted

Recent approaches have demonstrated the effectiveness of Vision Transformer (ViT) with attention mechanisms for domain generalization of Face Anti-Spoofing (FAS). However, current attention algorithms highlight all the salient objects (e.g., background objects, hair, glasses), which results in the f…

Cited by 0SourceScholar
2023

Cross-Domain Facial Expression Recognition via Disentangling Identity Representation

IJCAI 2023poster

Most existing cross-domain facial expression recognition (FER) works require target domain data to assist the model in analyzing distribution shifts to overcome negative effects. However, it is often hard to obtain expression images of the target domain in practical applications. Moreover, existing…

Cited by 9SourcePDFScholar
2023

Gloss-Free Sign Language Translation: Improving from Visual-Language Pretraining

ICCV 2023poster

Sign Language Translation (SLT) is a challenging task due to its cross-domain nature, involving the translation of visual-gestural language to text. Many previous methods employ an intermediate representation,i.e., gloss sequences, to facilitate SLT, thus transforming it into a two-stage task of sig…

Cited by 62PDFcodeScholar
2023

Mixture Uniform Distribution Modeling and Asymmetric Mix Distillation for Class Incremental Learning

AAAI 2023technical

Exemplar rehearsal-based methods with knowledge distillation (KD) have been widely used in class incremental learning (CIL) scenarios. However, they still suffer from performance degradation because of severely distribution discrepancy between training and test set caused by the limited storage memo…

Cited by 11SourcePDFScholar
2022

Decoupling and Recoupling Spatiotemporal Representation for RGB-D-Based Motion Recognition

CVPR 2022poster

Decoupling spatiotemporal representation refers to decomposing the spatial and temporal features into dimension-independent factors. Although previous RGB-D-based motion recognition methods have achieved promising performance through the tightly coupled multi-modal spatiotemporal representation, the…

Cited by 46PDFcodeScholar
2022

Nested Collaborative Learning for Long-Tailed Visual Recognition

CVPR 2022poster

The networks trained on the long-tailed dataset vary remarkably, despite the same training settings, which shows the great uncertainty in long-tailed learning. To alleviate the uncertainty, we propose a Nested Collaborative Learning (NCL), which tackles the problem by collaboratively learning multip…

Cited by 121PDFcodeScholar
2021

Regional Attention with Architecture-Rebuilt 3D Network for RGB-D Gesture Recognition

AAAI 2021technical

Human gesture recognition has drawn much attention in the area of computer vision. However, the performance of gesture recognition is always influenced by some gesture-irrelevant factors like the background and the clothes of performers. Therefore, focusing on the regions of hand/arm is important to…

2019

A Dataset and Benchmark for Large-Scale Multi-Modal Face Anti-Spoofing

CVPR 2019poster

Face anti-spoofing is essential to prevent face recognition systems from a security breach. Much of the progresses have been made by the availability of face anti-spoofing benchmark datasets in recent years. However, existing face anti-spoofing benchmarks have limited number of subjects (<=170) and…

Cited by 215PDFScholar