ECCV 2024poster2 citations
Multi-Modal Video Dialog State Tracking in the Wild
Adnen Abdessaied*, Lei Shi, Andreas Bulling
Abstract
"We present figures/mixeri con.pdf −−anovelvideodialogmodeloperatingoveragenericmulti− modalstatetrackingscheme.Currentmodelsthatclaimtoperf ormmulti−modalstatetrackingf allshortintwoma (1)T heyeithertrackonlyonemodality(mostlythevisualinput)or(2)theytargetsyntheticdatasetsthatdonotref lec worldin−the−wildscenarios.Ourmodeladdressesthesetwolimitationsinanattempttoclosethiscrucialresearch modalgraphstructurelearningmethod.Subsequently, thelearnedlocalgraphsandf eaturesareparsedtogethertof grainedgraphnodef eaturesareusedtoenhancethehiddenstatesof thebackboneV ision− LanguageM odel(V LM ). achievesnewstate−of −the−artresultsonfivechallengingbenchmarks."
BibTeX
@inproceedings{eccv2024_multimodalvideod,
title = {Multi-Modal Video Dialog State Tracking in the Wild},
author = {Adnen Abdessaied* and Lei Shi and Andreas Bulling},
booktitle = {ECCV 2024},
year = {2024}
}