← Search

Tran Minh Quan

2 accepted papers

2025

Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation

NeurIPS 2025poster

Vision-Language Foundation Models (VLMs), trained on large-scale multimodal datasets, have driven significant advances in Artificial Intelligence (AI) by enabling rich cross-modal reasoning. Despite their success in general domains, applying these models to medical imaging remains challenging due to…

Cited by 0SourceScholar
2021

ColorRL: Reinforced Coloring for End-to-End Instance Segmentation

CVPR 2021poster

Instance segmentation, the task of identifying and separating each individual object of interest in the image, is one of the actively studied research topics in computer vision. Although many feed-forward networks produce high-quality binary segmentation on different types of images, their final res…

Cited by 4PDFcodeScholar