← Search

Adnen Abdessaied

4 accepted papers

2025

V^2Dial: Unification of Video and Visual Dialog via Multimodal Experts

CVPR 2025poster

We present V2Dial - a novel expert-based model specifically geared towards simultaneously handling image and video input data for multimodal conversational tasks. Current multimodal models primarily focus on simpler tasks (e.g., VQA, VideoQA, video-text retrieval) and often neglect the more challeng…

Cited by 0SourcePDFScholar
2024

Limits of Theory of Mind Modelling in Dialogue-Based Collaborative Plan Acquisition

ACL 2024long

Recent work on dialogue-based collaborative plan acquisition (CPA) has suggested that Theory of Mind (ToM) modelling can improve missing knowledge prediction in settings with asymmetric skill-sets and knowledge. Although ToM was claimed to be important for effective collaboration, its real impact on…

Cited by 6SourcePDFScholar
2024

OLViT: Multi-Modal State Tracking via Attention-Based Embeddings for Video-Grounded Dialog

COLING 2024main

We present the Object Language Video Transformer (OLViT) – a novel model for video dialog operating over a multi-modal attention-based dialog state tracker. Existing video dialog models struggle with questions requiring both spatial and temporal localization within videos, long-term temporal reasoni…

Cited by 2SourcePDFScholar