← Search

Alvaro Soto

14 accepted papers

2026

CURE: Curriculum-guided Multi-task Training for Reliable Anatomy Grounded Report Generation

CVPR 2026

Medical vision-language models can automate the generation of radiology reports but struggle with accurate visual grounding and factual consistency. Existing models often misalign textual findings with visual evidence, leading to unreliable or weakly grounded predictions. We present "CURE", an error

Cited by 1SourcecodeScholar
2025

Data Distributional Properties As Inductive Bias for Systematic Generalization

CVPR 2025poster

Deep neural networks (DNNs) struggle at systematic generalization (SG). Several studies have evaluated the possibility of promoting SG through the proposal of novel architectures, loss functions, or training methodologies. Few studies, however, have focused on the role of training data properties in…

2025

Tracr-Injection: Distilling Algorithms into Pre-trained Language Models

ACL 2025finding

Motivated by the surge of large language models, there has been a push to formally characterize the symbolic abilities intrinsic to the transformer architecture. A programming language, called RASP, has been proposed, which can be directly compiled into transformer weights to implement these algorit…

2024

Extracting and Encoding: Leveraging Large Language Models and Medical Knowledge to Enhance Radiological Text Representation

ACL 2024findings

Advancing representation learning in specialized fields like medicine remains challenging due to the scarcity of expert annotations for text and images. To tackle this issue, we present a novel two-stage framework designed to extract high-quality factual statements from free-text radiology reports i…

2023

A Memory Model for Question Answering from Streaming Data Supported by Rehearsal and Anticipation of Coreference Information

ACL 2023findings

Existing question answering methods often assume that the input content (e.g., documents or videos) is always accessible to solve the task. Alternatively, memory networks were introduced to mimic the human process of incremental comprehension and compression of the information in a fixed-capacity me…

Cited by 6SourcePDFScholar
2023

PIVOT: Prompting for Video Continual Learning

CVPR 2023poster

Modern machine learning pipelines are limited due to data availability, storage quotas, privacy regulations, and expensive annotation processes. These constraints make it difficult or impossible to train and update large-scale models on such dynamic annotated sets. Continual learning directly approa…

Cited by 60SourcePDFScholar
2022

Bridging the Visual Semantic Gap in VLN via Semantically Richer Instructions

ECCV 2022poster

"The Visual-and-Language Navigation (VLN) task requires understanding a textual instruction to navigate a natural indoor environment using only visual information. While this is a trivial task for most humans, it is still an open problem for AI models. In this work, we hypothesize that poor use of t…

2021

Augmenting BERT-style Models with Predictive Coding to Improve Discourse-level Representations

EMNLP 2021main

Current language models are usually trained using a self-supervised scheme, where the main focus is learning representations at the word or sentence level. However, there has been limited progress in generating useful discourse-level representations. In this work, we propose to use ideas from predic…

Cited by 9SourcePDFScholar
2021

Optimizing Reusable Knowledge for Continual Learning via Metalearning

NeurIPS 2021poster

When learning tasks over time, artificial neural networks suffer from a problem known as Catastrophic Forgetting (CF). This happens when the weights of a network are overwritten during the training of a new task causing forgetting of old information. To address this issue, we propose MetA Reusable K…

2019

A Behavioral Approach to Visual Navigation with Graph Localization Networks

RSS 2019poster

Inspired by research in psychology, we introduce a behavioral approach for visual navigation using topological maps. Our goal is to enable a robot to navigate from one location to another, relying only on its visual observations and the topological map of the environment. To this end, we propose usi…

Cited by 123SourcePDFScholar
2018

End-to-End Joint Semantic Segmentation of Actors and Actions in Video

ECCV 2018poster

Traditional video understanding tasks include human action recognition and actor/object semantic segmentation. However, the combined task of providing semantic segmentation for different actor classes simultaneously with their action class remains a challenging but necessary task for many applicatio…

Cited by 52SourcePDFScholar
2016

A Hierarchical Pose-Based Approach to Complex Action Understanding Using Dictionaries of Actionlets and Motion Poselets

CVPR 2016poster

In this paper, we introduce a new hierarchical model for human action recognition that is able to categorize complex actions performed in videos. Our model is also able to perform spatio-temporal annotation of the atomic actions that compose the overall complex action. That is, for each atomic actio…

Cited by 66PDFScholar