← Search

Xiaohan Xing

10 accepted papers

2026

A Survey of Artificial Intelligence in Endoscopic Surgery Workflow: From Perception to Surgical Support

IJCAI 2026

Endoscopic surgery demands continuous real-time visual decision-making under severe constraints, including a limited field of view, motion blur, and dynamically deforming anatomy. These factors impose substantial cognitive load on surgeons and motivate the integration of artificial intelligence (AI)

Cited by 0Scholar
2026

H2-Surv: Hierarchical Hyperbolic Multimodal Representation Learning for Survival Prediction

CVPR 2026

Cancer survival prediction through multimodal learning that combines histopathology images with genomic data represents a promising research direction. However, current approaches still suffer from two key limitations. First, most methods operate in a Euclidean feature space, which makes it difficul

Cited by 0SourceScholar
2025

Asclepius: A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language Models

ACL 2025long

The significant breakthroughs of Medical Multi-Modal Large Language Models (Med-MLLMs) renovate modern healthcare with robust information synthesis and medical decision support. However, these models are often evaluated on benchmarks that are unsuitable for the Med-MLLMs due to the intricate nature…

2025

OccMamba: Semantic Occupancy Prediction with State Space Models

CVPR 2025poster

Training deep learning models for semantic occupancy prediction is challenging due to factors such as a large number of occupancy cells, severe occlusion, limited visual cues, complicated driving scenarios, etc. Recent methods often adopt transformer-based architectures given their strong capability…

2025

One Polyp Identifies All: One-Shot Polyp Segmentation with SAM via Cascaded Priors and Iterative Prompt Evolution

ICCV 2025poster

Polyp segmentation is vital for early colorectal cancer detection, yet traditional fully supervised methods struggle with morphological variability and domain shifts, requiring frequent retraining. Additionally, reliance on large-scale annotations is a major bottleneck due to the time-consuming and…

2025

WSI-LLaVA: A Multimodal Large Language Model for Whole Slide Image

ICCV 2025poster

Recent advances in computational pathology have introduced whole slide image (WSI)-level multimodal large language models (MLLMs) for automated pathological analysis. However, current WSI-level MLLMs face two critical challenges: limited explainability in their decision-making process and insufficie…

Cited by 0SourcePDFScholar
2023

Locate before Segment: Topology-guided Retinal Layer Segmentation in Optical Coherence Tomography Images

ICRA 2023poster

Optical Coherence Tomography (OCT) is a non-invasive imaging technique that is instrumental in retinal disease diagnosis and treatment. Segmentation of retinal layers in OCT is an essential step, but remains challenging for common pixel-wise segmentation methods usually fail to obtain the correct la…

Cited by 0SourceScholar
2022

Multiple Consistency Supervision based Semi-supervised OCT Segmentation using Very Limited Annotations

ICRA 2022poster

Optical Coherence Tomography (OCT) is a rapidly growing and promising imaging technique, enabling non-invasive high-resolution visualization of biological tissues. Segmentation of tissue structures from OCT scans is essen-tial for disease diagnosis but remains challenging for the blurry boundaries a…

Cited by 3SourceScholar
2021

Multibranch Learning for Angiodysplasia Segmentation with Attention-Guided Networks and Domain Adaptation

ICRA 2021poster

As a common cause of anemia and gastrointestinal bleeding, angiodysplasia (AD) diagnosis in wireless capsule endoscopy (WCE) images is important in clinical. Current manual review requires undivided concentration of the gastroenterologists, which is laborious and time-consuming. The development of c…

Cited by 2SourceScholar
2020

Diagnose like a Clinician: Third-order Attention Guided Lesion Amplification Network for WCE Image Classification

IROS 2020poster

Wireless capsule endoscopy (WCE) is a novel imaging tool that allows the noninvasive visualization of the entire gastrointestinal (GI) tract without causing discomfort to the patients. Although convolutional neural networks (CNNs) have obtained promising performance for the automatic lesion recognit…

Cited by 2SourceScholar