← Search

Andrés Villa

5 accepted papers

2026

CURE: Curriculum-guided Multi-task Training for Reliable Anatomy Grounded Report Generation

CVPR 2026

Medical vision-language models can automate the generation of radiology reports but struggle with accurate visual grounding and factual consistency. Existing models often misalign textual findings with visual evidence, leading to unreliable or weakly grounded predictions. We present "CURE", an error

Cited by 1SourcecodeScholar
2026

MoDA: Modulation Adapter for Fine-Grained Visual Understanding in Instructional MLLMs

ICML 2026poster

Multimodal Large Language Models (MLLMs) have achieved remarkable success in instruction-following tasks by integrating pretrained visual encoders with large language models (LLMs). However, existing approaches often struggle with fine-grained visual grounding due to semantic entanglement in visual …

Cited by 0SourceScholar
2023

PIVOT: Prompting for Video Continual Learning

CVPR 2023poster

Modern machine learning pipelines are limited due to data availability, storage quotas, privacy regulations, and expensive annotation processes. These constraints make it difficult or impossible to train and update large-scale models on such dynamic annotated sets. Continual learning directly approa…

Cited by 60SourcePDFScholar
2022

vCLIMB: A Novel Video Class Incremental Learning Benchmark

CVPR 2022oral

Continual learning (CL) is under-explored in the video domain. The few existing works contain splits with imbalanced class distributions over the tasks, or study the problem in unsuitable datasets. We introduce vCLIMB, a novel video continual learning benchmark. vCLIMB is a standardized test-bed to…

Cited by 46PDFScholar
2021

Augmenting BERT-style Models with Predictive Coding to Improve Discourse-level Representations

EMNLP 2021main

Current language models are usually trained using a self-supervised scheme, where the main focus is learning representations at the word or sentence level. However, there has been limited progress in generating useful discourse-level representations. In this work, we propose to use ideas from predic…

Cited by 9SourcePDFScholar