← Search

Juan Leon Alcazar

3 accepted papers

2026

CURE: Curriculum-guided Multi-task Training for Reliable Anatomy Grounded Report Generation

CVPR 2026

Medical vision-language models can automate the generation of radiology reports but struggle with accurate visual grounding and factual consistency. Existing models often misalign textual findings with visual evidence, leading to unreliable or weakly grounded predictions. We present "CURE", an error

Cited by 1SourcecodeScholar
2026

MoDA: Modulation Adapter for Fine-Grained Visual Understanding in Instructional MLLMs

ICML 2026poster

Multimodal Large Language Models (MLLMs) have achieved remarkable success in instruction-following tasks by integrating pretrained visual encoders with large language models (LLMs). However, existing approaches often struggle with fine-grained visual grounding due to semantic entanglement in visual …

Cited by 0SourceScholar
2020

Active Speakers in Context

CVPR 2020poster

Current methods for active speaker detection focus on modeling audiovisual information from a single speaker. This strategy can be adequate for addressing single-speaker scenarios, but it prevents accurate detection when the task is to identify who of many candidate speakers are talking. This paper…

Cited by 104PDFcodeScholar