← Search

Lanfen Lin

20 accepted papers

2026

CG-DMER: Hybrid Contrastive-Generative Framework for Disentangled Multimodal ECG Representation Learning

ICASSP 2026oral

Accurate interpretation of electrocardiogram (ECG) signals is crucial for diagnosing cardiovascular diseases. Recent multimodal approaches that integrate ECGs with accompanying clinical reports show strong potential, but they still face two main concerns from a modality perspective: (1) intra-modali…

Cited by 0SourcePDFScholar
2026

Dynamic Summary Generation for Interpretable Multimodal Depression Detection

ICASSP 2026poster

Depression remains widely underdiagnosed and undertreated because stigma and subjective symptom ratings hinder reliable screening. To address this challenge, we propose a coarse-to-fine, multi-stage framework that leverages large language models (LLMs) for accurate and interpretable detection. The p…

Cited by 0SourcePDFScholar
2026

Language-guided Frequency Modulation for Large Vision-Language Models

CVPR 2026

Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in visual reasoning across diverse tasks. These tasks place different demands on visual representations: some prioritize high-level global context, while others emphasize fine-grained local details. However, most existing

Cited by 0SourceScholar
2026

Taming the Phantom: Token-Asymmetric Filtering for Hallucination Mitigation in Large Vision-Language Models

AAAI 2026technical

Hallucination in Large Vision-Language Models (LVLMs) remains a critical challenge, undermining their reliability in real-world applications. Existing studies have investigated the causes of hallucination at the modality level and proposed effective strategies. However, interaction patterns beyond

Cited by 0SourcePDFScholar
2025

Enhanced Multimodal Depression Detection With Emotion Prompts

ICASSP 2025accepted

Depression is a pervasive mental health disorder that remains frequently undiagnosed and untreated due to societal barriers and the subjective nature of its symptoms. Leveraging recent advances in large language models (LLMs), we propose a novel depression detection pipeline that generates emotion p…

Cited by 0SourceScholar
2025

M2OST: Many-to-one Regression for Predicting Spatial Transcriptomics from Digital Pathology Images

AAAI 2025technical

The advancement of Spatial Transcriptomics (ST) has facilitated the spatially-aware profiling of gene expressions based on histopathology images. Although ST data offers valuable insights into the micro-environment of tumors, its acquisition cost remains expensive. Therefore, directly predicting the…

2025

Region-aware Anchoring Mechanism for Efficient Referring Visual Grounding

ICCV 2025poster

Referring Visual Grounding (RVG) tasks revolve around utilizing vision-language interactions to incorporate object information from language expressions, thereby enabling targeted object detection or segmentation within images. Transformer-based methods have enabled effective interaction through att…

Cited by 0SourcePDFScholar
2024

Combinatorial CNN-Transformer Learning with Manifold Constraints for Semi-supervised Medical Image Segmentation

AAAI 2024technical

Semi-supervised learning (SSL), as one of the dominant methods, aims at leveraging the unlabeled data to deal with the annotation dilemma of supervised learning, which has attracted much attentions in the medical image segmentation. Most of the existing approaches leverage a unitary network by conv…

Cited by 7SourcePDFScholar
2024

Going Beyond Multi-Task Dense Prediction with Synergy Embedding Models

CVPR 2024poster

Multi-task visual scene understanding aims to leverage the relationships among a set of correlated tasks which are solved simultaneously by embedding them within a uni- fied network. However most existing methods give rise to two primary concerns from a task-level perspective: (1) the lack of task-i…

Cited by 5SourcePDFScholar
2024

IRLSG: Invariant Representation Learning for Single-Domain Generalization in Medical Image Segmentation

ICASSP 2024accepted

Single-domain generalization (SDG) can efficiently enhance model generalization while avoiding high annotation costs and privacy concerns. However, existing SDG methods are mainly based on data manipulation and meta-learning, which are not efficient enough due to the limited generalization performan…

Cited by 0SourceScholar
2023

ClassFormer: Exploring Class-Aware Dependency with Transformer for Medical Image Segmentation

AAAI 2023technical

Vision Transformers have recently shown impressive performances on medical image segmentation. Despite their strong capability of modeling long-range dependencies, the current methods still give rise to two main concerns in a class-level perspective: (1) intra-class problem: the existing methods lac…

Cited by 6SourcePDFScholar
2023

HAP: Structure-Aware Masked Image Modeling for Human-Centric Perception

NeurIPS 2023poster

Model pre-training is essential in human-centric perception. In this paper, we first introduce masked image modeling (MIM) as a pre-training approach for this task. Upon revisiting the MIM training strategy, we reveal that human structure priors offer significant potential. Motivated by this insight…

2023

MCKD: Mutually Collaborative Knowledge Distillation For Federated Domain Adaptation And Generalization

ICASSP 2023accepted

Conventional unsupervised domain adaptation (UDA) and domain generalization (DG) methods rely on the assumption that all source domains can be directly accessed and combined for model training. However, this centralized training strategy may violate privacy policies in many real-world applications.…

Cited by 0SourceScholar
2023

SLViT: Scale-Wise Language-Guided Vision Transformer for Referring Image Segmentation

IJCAI 2023poster

Referring image segmentation aims to segment an object out of an image via a specific language expression. The main concept is establishing global visual-linguistic relationships to locate the object and identify boundaries using details of the image. Recently, various Transformer-based techniques h…

2023

SemiCVT: Semi-Supervised Convolutional Vision Transformer for Semantic Segmentation

CVPR 2023poster

Semi-supervised learning improves data efficiency of deep models by leveraging unlabeled samples to alleviate the reliance on a large set of labeled samples. These successes concentrate on the pixel-wise consistency by using convolutional neural networks (CNNs) but fail to address both global learni…

Cited by 25SourcePDFScholar
2022

Mixed Transformer U-Net for Medical Image Segmentation

ICASSP 2022accepted

Though U-Net has achieved tremendous success in medical image segmentation tasks, it lacks the ability to explicitly model long-range dependencies. Therefore, Vision Transformers have emerged as alternative segmentation structures recently, for their innate ability of capturing long-range correlatio…

Cited by 0SourceScholar
2022

ScaleFormer: Revisiting the Transformer-based Backbones from a Scale-wise 
Perspective for Medical Image Segmentation

IJCAI 2022poster

Recently, a variety of vision transformers have been developed as their capability of modeling long-range dependency. In current transformer-based backbones for medical image segmentation, convolutional layers were replaced with pure transformers, or transformers were added to the deepest encoder to…

2021

Graph-BAS3Net: Boundary-Aware Semi-Supervised Segmentation Network With Bilateral Graph Convolution

ICCV 2021poster

Semi-supervised learning (SSL) algorithms have attracted much attentions in medical image segmentation by leveraging unlabeled data, which challenge in acquiring massive pixel-wise annotated samples. However, most of the existing SSLs neglected the geometric shape constraint in object, leading to un…

Cited by 23PDFScholar
2021

Graph-Based Pyramid Global Context Reasoning With a Saliency- Aware Projection for Covid-19 Lung Infections Segmentation

ICASSP 2021accepted

Coronavirus Disease 2019 (COVID-19) has rapidly spread in 2020, emerging a mass of studies for lung infection segmentation from CT images. Though many methods have been proposed for this issue, it is a challenging task because of infections of various size appearing in different lobe zones. To tackl…

Cited by 0SourceScholar
2020

UNet 3+: A Full-Scale Connected UNet for Medical Image Segmentation

ICASSP 2020accepted

Recently, a growing interest has been seen in deep learning-based semantic segmentation. UNet, which is one of deep learning networks with an encoder-decoder architecture, is widely used in medical image segmentation. Combining multi-scale features is one of important factors for accurate segmentati…

Cited by 0SourceScholar