← Search

Xiaomeng Li

40 accepted papers

2026

A Cognitive Process-Inspired Architecture for Subject-Agnostic Brain Visual Decoding

ICLR 2026poster

Subject-agnostic brain decoding, which aims to reconstruct continuous visual experiences from fMRI without subject-specific training, holds great potential for clinical applications. However, this direction remains underexplored due to challenges in cross-subject generalization and the complex natur…

Cited by 0SourcecodeScholar
2026

AttTok: Marrying Attribute Tokens with Generative Pre-trained Vision-Language Models towards Medical Image Understanding

ICLR 2026poster

Recent generative pre-trained vision–language (GPTv) models have achieved remarkable success in multi-modal understanding, inspiring their adaptation to medical imaging tasks such as disease diagnosis and visual question answering (VQA). However, current instruction-tuned GPTv models suffer from two…

Cited by 0SourceScholar
2026

CaPro: Curvilinear-aware Prompt Learning with Single Unlabeled Image for Cost-effective Curvilinear Structure Segmentation

AAAI 2026technical

Curvilinear structure segmentation (CSS) plays a vital role in industrial applications, including medical imaging and structural health monitoring. Recently, the strong capacity of the Segment Anything Model (SAM) has inspired its downstream application in CSS tasks. To adapt SAM to CSS tasks, previ

Cited by 0SourcePDFScholar
2026

DenseMLLM: Standard Multimodal LLMs are Intrinsic Dense Predictors

ICML 2026poster

Multimodal Large Language Models (MLLMs) have demonstrated exceptional capabilities in high-level visual understanding. However, extending these models to fine-grained dense prediction tasks, such as semantic segmentation and depth estimation, typically necessitates the incorporation of complex, tas…

Cited by 0SourceScholar
2026

Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression

ICML 2026poster

Multimodal large language models (MLLMs) struggle with numerical regression under longtailed target distributions. Token-level supervised fine-tuning (SFT) and point-wise regression rewards bias learning toward high-density regions, leading to regression-to-the-mean behavior and poor tail performanc…

Cited by 0SourceScholar
2026

OFL-SAM2: Prompt SAM2 with Online Few-shot Learner for Efficient Medical Image Segmentation

AAAI 2026technical

The Segment Anything Model 2 (SAM2) has demonstrated remarkable promptable visual segmentation capabilities in video data, showing potential for extension to medical image segmentation (MIS) tasks involving 3D volumes and temporally correlated 2D image sequences. However, adapting SAM2 to MIS presen

Cited by 0SourcePDFScholar
2026

Prompt Yourself: Awakening Textual Semantics in 1D Visual Tokenizers

CVPR 2026

One-dimensional (1D) visual tokenizers offer notable semantic compactness by discarding local spatial priors, and have become increasingly popular for image reconstruction and generation tasks. However, such global and sequential representations struggle to preserve fine-grained visual content; simp

Cited by 0SourceScholar
2026

Scaling and Transferability of Annealing Strategies in Large Language Model Training

AAAI 2026technical

Learning rate scheduling is crucial for training large language models, yet understanding the optimal annealing strategies across different model configurations remains challenging. In this work, we investigate the transferability of annealing dynamics in large language model training and refine a g

Cited by 0SourcePDFScholar
2026

TLMA: Mitigating the Impact of Weakly Labeled Information for Video Anomaly Detection

CVPR 2026

Weakly Supervised Video Anomaly Detection (WSVAD) aims to localize abnormal segments using only video-level labels during training.Although the paradigm significantly reduces annotation costs, the coarse-grained labels fail to precisely describe the full videos, resulting in the introduction of subs

Cited by 0SourceScholar
2025

GDFDNet: A Novel Graph-Based Dynamically Fused Dual-Stream Network for Accuracy Prohibited Items Detection

ICASSP 2025accepted

In various security inspection scenarios, prohibited items detection in X-ray images is of great significance for safeguarding public safety and effectively reducing the potential risks of crimes and terrorist activities. However, existing detection methods still face the challenge of overlapping im…

Cited by 0SourceScholar
2025

Gate to the Vessel: Residual Experts Restore What SAM Overlooks

NeurIPS 2025poster

Foundation segmentation models like Segment Anything (SAM) exhibit strong generalization on natural images but struggle with localized failures in medical imaging, especially on fine-grained structures such as vessels with complex morphology and indistinct boundaries. To address this, we propose Fin…

Cited by 0SourceScholar
2025

Interpretable Bilingual Multimodal Large Language Model for Diverse Biomedical Tasks

ICLR 2025poster

Several medical Multimodal Large Languange Models (MLLMs) have been developed to address tasks involving visual images with textual instructions across various medical modalities, achieving impressive results. Most current medical generalist models are region-agnostic, treating the entire image as…

Cited by 3SourcePDFScholar
2025

Leveraging Anatomical Consistency for Multi-Object Detection in Ultrasound Images via Source-free Unsupervised Domain Adaptation

AAAI 2025technical

Source-free unsupervised domain adaptation aims to eliminate domain shifts when data from the source domain and annotation from the target domain are not available. The multi-object detection tasks in medical image analysis are constrained by patient privacy and extremely huge annotation consumption…

2025

MuTri: Multi-view Tri-alignment for OCT to OCTA 3D Image Translation

CVPR 2025poster

Optical coherence tomography angiography (OCTA) shows its great importance in imaging microvascular networks by providing accurate 3D imaging of blood vessels, but it relies upon specialized sensors and expensive devices. For this reason, previous works show the potential to translate the readily av…

2025

Neurons: Emulating the Human Visual Cortex Improves Fidelity and Interpretability in fMRI-to-Video Reconstruction

ICCV 2025poster

Decoding visual stimuli from neural activity is essential for understanding the human brain. While fMRI methods have successfully reconstructed static images, fMRI-tovideo reconstruction faces challenges due to the need for capturing spatiotemporal dynamics like motion and scene transitions. Recent…

2025

PlantPCC: Dual Sampling and Multi-level Geometry-aware Contrastive Regularization for Plant Point Cloud Completion

ICASSP 2025accepted

Plant point cloud completion is essential for tasks like segmentation and surface reconstruction in plant phenotyping. Unlike the relatively simpler Computer-Aided Design models found in datasets like ShapeNet, plant point clouds are characterized by their rich geometric shapes and intricate edge fe…

Cited by 0SourceScholar
2025

Token Activation Map to Visually Explain Multimodal LLMs

ICCV 2025poster

Multimodal large language models (MLLMs) are broadly empowering various fields. Despite their advancements, the explainability of MLLMs remains less explored, hindering deeper understanding, model credibility, and effective visualization. Unlike conventional vision models (e.g., CNNs, ViTs, CLIP) th…

2025

ZEBRA: Towards Zero-Shot Cross-Subject Generalization for Universal Brain Visual Decoding

NeurIPS 2025poster

Recent advances in neural decoding have enabled the reconstruction of visual experiences from brain activity, positioning fMRI-to-image reconstruction as a promising bridge between neuroscience and computer vision. However, current methods predominantly rely on subject-specific models or require subj…

Cited by 0SourcecodeScholar
2024

AllSpark: Reborn Labeled Features from Unlabeled in Transformer for Semi-Supervised Semantic Segmentation

CVPR 2024poster

Semi-supervised semantic segmentation (SSSS) has been proposed to alleviate the burden of time-consuming pixel-level manual labeling which leverages limited labeled data along with larger amounts of unlabeled data. Current state-of-the-art methods train the labeled data with ground truths and unlabe…

2024

Bidirectional Recurrence for Cardiac Motion Tracking with Gaussian Process Latent Coding

NeurIPS 2024poster

Quantitative analysis of cardiac motion is crucial for assessing cardiac function. This analysis typically uses imaging modalities such as MRI and Echocardiograms that capture detailed image sequences throughout the heartbeat cycle. Previous methods predominantly focused on the analysis of image pai…

2024

C^2RV: Cross-Regional and Cross-View Learning for Sparse-View CBCT Reconstruction

CVPR 2024poster

Cone beam computed tomography (CBCT) is an important imaging technology widely used in medical scenarios such as diagnosis and preoperative planning. Using fewer projection views to reconstruct CT also known as sparse-view reconstruction can reduce ionizing radiation and further benefit intervention…

2024

Energy-Based Concept Bottleneck Models: Unifying Prediction, Concept Intervention, and Probabilistic Interpretations

ICLR 2024poster

Existing methods, such as concept bottleneck models (CBMs), have been successful in providing concept-based interpretations for black-box deep learning models. They typically work by predicting concepts given the input and then predicting the final class label given the predicted concepts. However,…

2024

Holistic Autonomous Driving Understanding by Bird's-Eye-View Injected Multi-Modal Large Models

CVPR 2024poster

The rise of multimodal large language models (MLLMs) has spurred interest in language-based driving tasks. However existing research typically focuses on limited tasks and often omits key multi-view and temporal information which is crucial for robust autonomous driving. To bridge these gaps we intr…

2024

M3-UDA: A New Benchmark for Unsupervised Domain Adaptive Fetal Cardiac Structure Detection

CVPR 2024poster

The anatomical structure detection of fetal cardiac views is crucial for diagnosing fetal congenital heart disease. In practice there is a large domain gap between different hospitals' data such as the variable data quality due to differences in acquisition equipment. In addition accurate annotation…

2024

Multimodal Survival Ensemble Network: Integrating Genomic and Histopathological Insights for Enhanced Cancer Prognosis

ICASSP 2024accepted

Cancer’s inherent heterogeneity demands a multimodal approach to provide an accurate prognosis, taking into account histological, clinical, and genomic data. As the field of artificial intelligence evolves with advancements in multimodal learning, its role in survival analysis becomes increasingly c…

Cited by 0SourceScholar
2024

Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?

NeurIPS 2024poster

How can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified metrics, unfair comparisons, and short-term outcome pressure. As a consequence, good performance on standard benchmarks…

2024

Unsupervised Domain Adaptation for Anatomical Structure Detection in Ultrasound Images

ICML 2024poster

Models trained on ultrasound images from one institution typically experience a decline in effectiveness when transferred directly to other institutions. Moreover, unlike natural images, dense and overlapped structures exist in fetus ultrasound images, making the detection of structures more challen…

Cited by 7SourcePDFScholar
2023

GraphEcho: Graph-Driven Unsupervised Domain Adaptation for Echocardiogram Video Segmentation

ICCV 2023oral

Echocardiogram video segmentation plays an important role in cardiac disease diagnosis. This paper studies the unsupervised domain adaption (UDA) for echocardiogram video segmentation, where the goal is to generalize the model trained on the source domain to other unlabeled target domains. Existing…

Cited by 17PDFcodeScholar
2023

Semi-Supervised Contrastive Learning for Deep Regression with Ordinal Rankings from Spectral Seriation

NeurIPS 2023poster

Contrastive learning methods can be applied to deep regression by enforcing label distance relationships in feature space. However, these methods are limited to labeled data only unlike for classification, where unlabeled data can be used for contrastive pretraining. In this work, we extend contrast…

2023

Semi-Supervised Deep Regression with Uncertainty Consistency and Variational Model Ensembling via Bayesian Neural Networks

AAAI 2023technical

Deep regression is an important problem with numerous applications. These range from computer vision tasks such as age estimation from photographs, to medical tasks such as ejection fraction estimation from echocardiograms for disease tracking. Semi-supervised approaches for deep regression are nota…

2023

Towards Distribution-Agnostic Generalized Category Discovery

NeurIPS 2023poster

Data imbalance and open-ended distribution are two intrinsic characteristics of the real visual world. Though encouraging progress has been made in tackling each challenge separately, few works dedicated to combining them towards real-world scenarios. While several previous works have focused on cla…

2023

Towards Generic Semi-Supervised Framework for Volumetric Medical Image Segmentation

NeurIPS 2023poster

Volume-wise labeling in 3D medical images is a time-consuming task that requires expertise. As a result, there is growing interest in using semi-supervised learning (SSL) techniques to train models with limited labeled data. However, the challenges and practical applications extend beyond SSL to set…

2022

Enhancing Pseudo Label Quality for Semi-supervised Domain-Generalized Medical Image Segmentation

AAAI 2022technical

Generalizing the medical image segmentation algorithms to unseen domains is an important research topic for computer-aided diagnosis and surgery. Most existing methods require a fully labeled dataset in each source domain. Although some researchers developed a semi-supervised domain generalized meth…

2022

RSCFed: Random Sampling Consensus Federated Semi-Supervised Learning

CVPR 2022poster

Federated semi-supervised learning (FSSL) aims to derive a global model by jointly training fully-labeled and fully-unlabeled clients. The existing approaches work well when local clients have independent and identically distributed (IID) data but fail to generalize to a more practical FSSL setting,…

Cited by 90PDFcodeScholar
2022

Separated Contrastive Learning for Organ-at-Risk and Gross-Tumor-Volume Segmentation with Limited Annotation

AAAI 2022technical

Automatic delineation of organ-at-risk (OAR) and gross-tumor-volume (GTV) is of great significance for radiotherapy planning. However, it is a challenging task to learn powerful representations for accurate delineation under limited pixel (voxel)-wise annotations. Contrastive learning at pixel-level…

2022

Uncertainty Estimation via Response Scaling for Pseudo-Mask Noise Mitigation in Weakly-Supervised Semantic Segmentation

AAAI 2022technical

Weakly-Supervised Semantic Segmentation (WSSS) segments objects without heavy burden of dense annotation. While as a price, generated pseudo-masks exist obvious noisy pixels, which result in sub-optimal segmentation models trained over these pseudo-masks. But rare studies notice or work on this prob…

2021

Support-Set Based Cross-Supervision for Video Grounding

ICCV 2021poster

Current approaches for video grounding propose kinds of complex architectures to capture the video-text relations, and have achieved impressive improvements. However, it is hard to learn the complicated multi-modal relations by only architecture designing in fact. In this paper, we introduce a novel…

Cited by 53PDFScholar