← Search

Manuel Hochmeister

1 accepted papers

2024

OLViT: Multi-Modal State Tracking via Attention-Based Embeddings for Video-Grounded Dialog

COLING 2024main

We present the Object Language Video Transformer (OLViT) – a novel model for video dialog operating over a multi-modal attention-based dialog state tracker. Existing video dialog models struggle with questions requiring both spatial and temporal localization within videos, long-term temporal reasoni…

Cited by 2SourcePDFScholar