← Search

Nicolas Padoy

13 accepted papers

2026

From Panel to Pixel: Zoom-In Vision-Language Pretraining from Biomedical Scientific Literature

CVPR 2026

There is growing interest in biomedical vision--language models trained on scientific literature. However, most pipelines compress rich multi-panel figures and long captions into coarse figure-level pairs, discarding the fine-grained correspondences clinicians rely on when zooming into local structu

Cited by 0SourceScholar
2026

Where It Moves, It Matters: Referring Surgical Instrument Segmentation via Motion

AAAI 2026technical

Enabling intuitive, language-driven interaction with surgical scenes is a critical step toward intelligent operating rooms and autonomous surgical robotic assistance. However, the task of referring segmentation, localizing surgical instruments based on natural language descriptions, remains underexp

Cited by 0SourcePDFScholar
2025

CholecTrack20: A Multi-Perspective Tracking Dataset for Surgical Tools

CVPR 2025poster

Tool tracking in surgical videos is essential for advancing computer-assisted interventions, such as skill assessment, safety zone estimation, and human-machine collaboration. However, the lack of context-rich datasets limits AI applications in this field. Existing datasets rely on overly generic tr…

2025

Learning from Synchronization: Self-Supervised Uncalibrated Multi-View Person Association in Challenging Scenes

CVPR 2025poster

Multi-view person association is a fundamental step towards multi-view analysis of human activities. Although the person re-identification features have been proven effective, they become unreliable in challenging scenes where persons share similar appearances. Therefore, cross-view geometric constr…

2025

Medical Multimodal Model Stealing Attacks via Adversarial Domain Alignment

AAAI 2025technical

Medical multimodal large language models (MLLMs) are becoming an instrumental part of healthcare systems, assisting medical personnel with decision making and results analysis. Models for radiology report generation are able to interpret medical imagery, thus reducing the workload of radiologists. A…

Cited by 1SourcePDFScholar
2025

OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining

ICCV 2025poster

Vision-language pretraining (VLP) enables open-world generalization beyond predefined labels, a critical capability in surgery due to the diversity of procedures, instruments, and patient anatomies. However, applying VLP to ophthalmic surgery presents unique challenges, including limited vision-lang…

2024

Procedure-Aware Surgical Video-language Pretraining with Hierarchical Knowledge Augmentation

NeurIPS 2024spotlight

Surgical video-language pretraining (VLP) faces unique challenges due to the knowledge domain gap and the scarcity of multi-modal data. This study aims to bridge the gap by addressing issues regarding textual information loss in surgical lecture videos and the spatial-temporal challenges of surgical…

2024

SelfPose3d: Self-Supervised Multi-Person Multi-View 3d Pose Estimation

CVPR 2024poster

We present a new self-supervised approach SelfPose3d for estimating 3d poses of multiple persons from multiple camera views. Unlike current state-of-the-art fully-supervised methods our approach does not require any 2d or 3d ground-truth poses and uses only the multi-view input images from a calibra…

2023

Why Is the Winner the Best?

CVPR 2023poster

International benchmarking competitions have become fundamental for the comparative performance assessment of image analysis methods. However, little attention has been given to investigating what can be learnt from these competitions. Do they really generate scientific progress? What are common and…

Cited by 29SourcePDFScholar
2021

A Kinematic Bottleneck Approach for Pose Regression of Flexible Surgical Instruments Directly From Images

RA-L 2021

3-D pose estimation of instruments is a crucial step towards automatic scene understanding in robotic minimally invasive surgery. Although robotic systems can potentially directly provide joint values, this information is not commonly exploited inside the operating room, due to its possible unreliab

Cited by 23SourceScholar
2019

Self-Supervised Surgical Tool Segmentation using Kinematic Information

ICRA 2019poster

Surgical tool segmentation in endoscopic images is the first step towards pose estimation and (sub-)task automation in challenging minimally invasive surgical operations. While many approaches in the literature have shown great results using modern machine learning methods such as convolutional neur…

Cited by 48SourceScholar
2017

Pose optimization of a C-arm imaging device to reduce intraoperative radiation exposure of staff and patient during interventional procedures

ICRA 2017poster

Minimally-invasive (MI) procedures are becoming more popular and frequent due to their benefits such as reduced patient trauma and hospitalization time. However, several common types of MI interventions are performed under X-ray guidance, which exposes both patients and staff to harmful ionizing rad…

Cited by 14SourceScholar