← Search

Yueming Jin

26 accepted papers

2026

3DMedAgent: Unified Perception-to-Understanding for 3D Medical Analysis

ICML 2026poster

3D CT analysis spans a continuum from low-level perception to high-level clinical understanding. Existing 3D-oriented analysis methods adopt either isolated task-specific modeling or task-agnostic end-to-end paradigms to produce one-hop outputs, impeding the systematic accumulation of perceptual evi…

Cited by 2SourceScholar
2026

Granulon: Awakening Pixel-Level Visual Encoders with Adaptive Multi-Granularity Semantics for MLLM

CVPR 2026

Recent advances in multimodal large language models largely rely on CLIP-based visual encoders, which emphasize global semantic alignment but struggle with fine-grained visual understanding. In contrast, DINOv3 provides strong pixel-level perception yet lacks coarse-grained semantic abstraction, lea

Cited by 0SourcecodeScholar
2026

MedAgent-Pro: Towards Evidence-based Multi-modal Medical Diagnosis via Reasoning Agentic Workflow

ICLR 2026poster

Modern clinical diagnosis relies on the comprehensive analysis of multi-modal patient data, drawing on medical expertise to ensure systematic and rigorous reasoning. Recent advances in Vision–Language Models (VLMs) and agent-based methods are reshaping medical diagnosis by effectively integrating mu…

Cited by 0SourcecodeScholar
2026

Tackling Dual-stage Missing Modalities in Brain Tumor Segmentation via Robust Modality Reconstruction and Prompt-guided Modality Adaptation

AAAI 2026technical

Addressing missing modalities is a critical challenge in multimodal brain tumor segmentation. Most existing approaches merely handle modality-incomplete inputs during inference, assuming a full set of modalities for all training samples. However, this unrealistic assumption limits the usage of abund

Cited by 0SourcePDFScholar
2026

Uncovering Hidden Triggers: Backdoor Attribution in Language Models

ICML 2026poster

Fine-tuned Large Language Models (LLMs) are vulnerable to backdoor attacks through data poisoning, yet the internal mechanisms governing these attacks remain a black box. Previous research on interpretability for LLM safety tends to focus on alignment, jailbreak, and hallucination, but overlooks bac…

Cited by 0SourceScholar
2026

Unleashing the Power of Image-Tabular Self-Supervised Learning via Breaking Cross-Tabular Barriers

AAAI 2026technical

Multi-modal learning integrating medical images and tabular data has significantly advanced clinical decision-making in recent years. Self-Supervised Learning (SSL) has emerged as a powerful paradigm for pretraining these models on large-scale unlabeled image-tabular data, aiming to learn discrimina

Cited by 0SourcePDFScholar
2025

Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools

ACL 2025long

We introduce Agentic Reasoning, a framework that enhances large language model (LLM) reasoning by integrating external tool-using agents. Agentic Reasoning dynamically leverages web search, code execution, and structured memory to address complex problems requiring deep research. A key innovation in…

Cited by 0SourcePDFScholar
2025

Automatically Identify and Rectify: Robust Deep Contrastive Multi-view Clustering in Noisy Scenarios

ICML 2025spotlight

Leveraging the powerful representation learning capabilities, deep multi-view clustering methods have demonstrated reliable performance by effectively integrating multi-source information from diverse views in recent years. Most existing methods rely on the assumption of clean views. However, noise…

2025

Generalized Deep Multi-view Clustering via Causal Learning with Partially Aligned Cross-view Correspondence

ICCV 2025poster

Multi-view clustering (MVC) aims to explore the common clustering structure across multiple views. Many existing MVC methods heavily rely on the assumption of view consistency, where alignments for corresponding samples across different views are ordered in advance. However, real-world scenarios oft…

Cited by 0SourcePDFScholar
2025

Medical Graph RAG: Evidence-based Medical Large Language Model via Graph Retrieval-Augmented Generation

ACL 2025long

We introduce MedGraphRAG, a novel graph-based Retrieval-Augmented Generation (RAG) framework designed to enhance LLMs in generating evidence-based medical responses, improving safety and reliability with private medical data. We introduce Triple Graph Construction and U-Retrieval to enhance GraphRAG…

2025

Multi-scale Temporal Prediction via Incremental Generation and Multi-agent Collaboration

NeurIPS 2025poster

Accurate temporal prediction is the bridge between comprehensive scene understanding and embodied artificial intelligence. However, predicting multiple fine-grained states of scene at multiple temporal scales is difficult for vision-language models. We formalize the Multi‐Scale Temporal Prediction (…

Cited by 0SourceScholar
2024

ANEDL: Adaptive Negative Evidential Deep Learning for Open-Set Semi-supervised Learning

AAAI 2024technical

Semi-supervised learning (SSL) methods assume that labeled data, unlabeled data and test data are from the same distribution. Open-set semi-supervised learning (Open-set SSL) con- siders a more practical scenario, where unlabeled data and test data contain new categories (outliers) not observed in l…

Cited by 5SourcePDFScholar
2024

LLM-Assisted Multi-Teacher Continual Learning for Visual Question Answering in Robotic Surgery

ICRA 2024poster

Visual question answering (VQA) can be fundamentally crucial for promoting robotic-assisted surgical education. In practice, the needs of trainees are constantly evolving, such as learning more surgical types and adapting to new surgical instruments/techniques. Therefore, continually updating the VQ…

Cited by 17SourcecodeScholar
2024

MedSegDiff-V2: Diffusion-Based Medical Image Segmentation with Transformer

AAAI 2024technical

The Diffusion Probabilistic Model (DPM) has recently gained popularity in the field of computer vision, thanks to its image generation applications, such as Imagen, Latent Diffusion Models, and Stable Diffusion, which have demonstrated impressive capabilities and sparked much discussion within the c…

2024

Think Step by Step: Chain-of-Gesture Prompting for Error Detection in Robotic Surgical Videos

RA-L 2024

Despite advancements in robotic systems and surgical data science, ensuring safe execution in robot-assisted minimally invasive surgery (RMIS) remains challenging. Current methods for surgical error detection typically involve two parts: identifying gestures and then detecting errors within each ges

Cited by 11SourcecodeScholar
2023

Dynamic Interactive Relation Capturing via Scene Graph Learning for Robotic Surgical Report Generation

ICRA 2023poster

For robot-assisted surgery, an accurate surgical report reflects clinical operations during surgery and helps document entry tasks, post-operative analysis and follow-up treatment. It is a challenging task due to many complex and diverse interactions between instruments and tissues in the surgical s…

Cited by 18SourceScholar
2023

Keep Your Eye on the Best: Contrastive Regression Transformer for Skill Assessment in Robotic Surgery

RA-L 2023

This letter proposes a novel video-based, contrastive regression architecture, Contra-Sformer, for automated surgical skill assessment in robot-assisted surgery. The proposed framework is structured to capture the differences in the surgical performance, between a test video and a reference video wh

Cited by 32SourceScholar
2022

Personalizing Federated Medical Image Segmentation via Local Calibration

ECCV 2022poster

"Medical image segmentation under federated learning (FL) is a promising direction by allowing multiple clinical sites to collaboratively learn a global model without centralizing datasets. However, using a single model to adapt to various data distributions from different sites is extremely challen…

2022

Pseudo-label Guided Cross-video Pixel Contrast for Robotic Surgical Scene Segmentation with Limited Annotations

IROS 2022poster

Surgical scene segmentation is fundamentally crucial for prompting cognitive assistance in robotic surgery. However, pixel-wise annotating surgical video in a frame-by-frame manner is expensive and time consuming. To greatly reduce the labeling burden, in this work, we study semi-supervised scene se…

Cited by 6SourcecodeScholar
2022

TraSeTR: Track-to-Segment Transformer with Contrastive Query for Instance-level Instrument Segmentation in Robotic Surgery

ICRA 2022poster

Surgical instrument segmentation - in general a pixel classification task - is fundamentally crucial for promoting cognitive intelligence in robot-assisted surgery (RAS). However, previous methods are struggling with discriminating instrument types and instances. To address above issues, we explore…

Cited by 49SourceScholar
2021

Accurate Grid Keypoint Learning for Efficient Video Prediction

IROS 2021poster

Video prediction methods generally consume substantial computing resources in training and deployment, among which keypoint-based approaches show promising improvement in efficiency by simplifying dense image prediction to light keypoint prediction. However, keypoint locations are often modeled only…

Cited by 18SourcecodeScholar
2021

Domain Adaptive Robotic Gesture Recognition with Unsupervised Kinematic-Visual Data Alignment

IROS 2021poster

Automated surgical gesture recognition is of great importance in robot-assisted minimally invasive surgery. However, existing methods assume that training and testing data are from the same domain, which suffers from severe performance degradation when a domain gap exists, such as the simulator and…

Cited by 4SourceScholar
2021

Modelling Neighbor Relation in Joint Space-Time Graph for Video Correspondence Learning

ICCV 2021poster

This paper presents a self-supervised method for learning reliable visual correspondence from unlabeled videos. We formulate the correspondence as finding paths in a joint space-time graph, where nodes are grid patches sampled from frames, and are linked by two type of edges: (i) neighbor relations…

Cited by 21PDFScholar
2021

One to Many: Adaptive Instrument Segmentation via Meta Learning and Dynamic Online Adaptation in Robotic Surgical Video

ICRA 2021poster

Surgical instrument segmentation in robot-assisted surgery (RAS) - especially that using learning-based models - relies on the assumption that training and testing videos are sampled from the same domain. However, it is impractical and expensive to collect and annotate sufficient data from every new…

Cited by 26SourceScholar
2021

Relational Graph Learning on Visual and Kinematics Embeddings for Accurate Gesture Recognition in Robotic Surgery

ICRA 2021poster

Automatic surgical gesture recognition is fundamentally important to enable intelligent cognitive assistance in robotic surgery. With recent advancement in robot-assisted minimally invasive surgery, rich information including surgical videos and robotic kinematics can be recorded, which provide comp…

Cited by 47SourceScholar
2020

Automatic Gesture Recognition in Robot-assisted Surgery with Reinforcement Learning and Tree Search

ICRA 2020poster

Automatic surgical gesture recognition is fundamental for improving intelligence in robot-assisted surgery, such as conducting complicated tasks of surgery surveillance and skill evaluation. However, current methods treat each frame individually and produce the outcomes without effective considerati…

Cited by 69SourceScholar