← Search

Jinkai Li

5 accepted papers

2026

Vision-Language Models Guided Graph Concept Reasoning for Interpretable Diabetic Retinopathy Diagnosis

AAAI 2026technical

Deep neural networks (DNNs) have significantly advanced diabetic retinopathy (DR) diagnosis, yet their black-box nature limits clinical acceptance due to a lack of interpretability. Concept bottleneck model (CBM) offers a promising solution by enabling concept-level reasoning and test-time intervent

Cited by 0SourcePDFScholar
2025

Component-wise Self-Correction Network for Human Motion Prediction

ICASSP 2025accepted

Human motion prediction is a fundamental task in human-robot interaction and self-driving. Many existing human motion prediction methods use one encoder to embed the historical human poses and one decoder to predict future motion poses. We believe that it is possible to estimate the deviation of the…

Cited by 0SourceScholar
2025

Deep Coarse-to-Fine Networks for Robust Segmentation and Pose Estimation of Surgical Suturing Threads

IROS 2025

Autonomous suturing is a critical challenge in robot-assisted surgery, where accurate segmentation and pose estimation of suturing threads are essential prerequisites. However, suturing threads are easily occluded by moving instruments and embedded in deformable tissues which make the task much more

Cited by 0SourceScholar
2024

Skill Learning in Robot-Assisted Micro-Manipulation Through Human Demonstrations with Attention Guidance

ICRA 2024poster

For the development of robotic systems for micromanipulation, it is challenging to design appropriate control strategies due to either the lack of sufficient information for feedback or the difficulty in extracting subtle yet critical visual features. With the same system under the teleoperated mode…

Cited by 2SourceScholar
2023

EasyGaze3D: Towards Effective and Flexible 3D Gaze Estimation from a Single RGB Camera

IROS 2023poster

Eye gaze can convey rich information of human intentions, which enables the social robots to comprehend the cognition and behavior of human targets. However, the existing 3D gaze estimation methods generally have high requirements either on the dedicated hardware or the quantity and quality of train…

Cited by 3SourceScholar