← Search

Wayner Barrios

3 accepted papers

2026

MoDA: Modulation Adapter for Fine-Grained Visual Understanding in Instructional MLLMs

ICML 2026poster

Multimodal Large Language Models (MLLMs) have achieved remarkable success in instruction-following tasks by integrating pretrained visual encoders with large language models (LLMs). However, existing approaches often struggle with fine-grained visual grounding due to semantic entanglement in visual …

Cited by 0SourceScholar
2023

Localizing Moments in Long Video Via Multimodal Guidance

ICCV 2023poster

The recent introduction of the large-scale, long-form MAD and Ego4D datasets has enabled researchers to investigate the performance of current state-of-the-art methods for video grounding in the long-form setup, with interesting findings: current grounding methods alone fail at tackling this challen…

Cited by 26PDFcodeScholar
2017

SCC: Semantic Context Cascade for Efficient Action Detection

CVPR 2017poster

Despite the recent advances in large-scale video analysis, action detection remains as one of the most challenging unsolved problems in computer vision. This snag is in part due to the large volume of data that needs to be analyzed to detect actions in videos. Existing approaches have mitigated the…

Cited by 111PDFScholar