← Search

Amos You

1 accepted papers

2024

TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering

EMNLP 2024main

Recently, image-based Large Multimodal Models (LMMs) have made significant progress in video question-answering (VideoQA) using a frame-wise approach by leveraging large-scale pretraining in a zero-shot manner. Nevertheless, these models need to be capable of finding relevant information, extracting…