← Search

Yongtao Hao

2 accepted papers

2025

KIA: Knowledge-Guided Implicit Vision-Language Alignment for Chest X-Ray Report Generation

COLING 2025main

Report generation (RG) faces challenges in understanding complex medical images and establishing cross-modal semantic alignment in radiology image-report pairs. Previous methods often overlook fine-grained cross-modal interaction, leading to insufficient understanding of detailed information. Recent…

2025

ROD-MLLM: Towards More Reliable Object Detection in Multimodal Large Language Models

CVPR 2025poster

Multimodal large language models (MLLMs) have demonstrated strong language understanding and generation capabilities, excelling in visual tasks like referring and grounding. However, due to task type limitations and dataset scarcity, existing MLLMs only ground objects present in images and cannot re…

Cited by 0SourcePDFScholar