← Search

Erwin M. Bakker

3 accepted papers

2023

COCA: COllaborative CAusal Regularization for Audio-Visual Question Answering

AAAI 2023technical

Audio-Visual Question Answering (AVQA) is a sophisticated QA task, which aims at answering textual questions over given video-audio pairs with comprehensive multimodal reasoning. Through detailed causal-graph analyses and careful inspections of their learning processes, we reveal that AVQA models ar…

Cited by 21SourcePDFScholar
2021

Lifelong Person Re-Identification via Adaptive Knowledge Accumulation

CVPR 2021poster

Person ReID methods always learn through a stationary domain that is fixed by the choice of a given dataset. In many contexts (e.g., lifelong learning), those methods are ineffective because the domain is continually changing in which case incremental learning over multiple domains is required poten…

Cited by 115PDFcodeScholar
2017

Learning a Recurrent Residual Fusion Network for Multimodal Matching

ICCV 2017poster

A major challenge in matching between vision and language is that they typically have completely different features and representations. In this work, we introduce a novel bridge between the modality-specific representations by creating a co-embedding space based on a recurrent residual fusion (RRF)…

Cited by 184PDFScholar