← Search

Ngoc Q. K. Duong

9 accepted papers

2021

On the Hidden Treasure of Dialog in Video Question Answering

ICCV 2021poster

High-level understanding of stories in video such as movies and TV shows from raw data is extremely challenging. Modern video question answering (VideoQA) systems often use additional human-made sources like plot synopses, scripts, video descriptions or knowledge bases. In this work, we present a ne…

Cited by 12PDFcodeScholar
2021

Self-Attention Generative Adversarial Network for Speech Enhancement

ICASSP 2021accepted

Existing generative adversarial networks (GANs) for speech enhancement solely rely on the convolution operation, which may obscure temporal dependencies across the sequence input. To remedy this issue, we propose a self-attention layer adapted from non-local attention, coupled with the convolutional…

Cited by 0SourceScholar
2019

VideoMem: Constructing, Analyzing, Predicting Short-Term and Long-Term Video Memorability

ICCV 2019poster

Humans share a strong tendency to memorize/forget some of the visual information they encounter. This paper focuses on understanding the intrinsic memorability of visual content. To address this challenge, we introduce a large scale dataset (VideoMem) composed of 10,000 videos with memorability scor…

Cited by 68PDFScholar
2018

Deep Learning for Predicting Image Memorability

ICASSP 2018accepted

Memorability of media content such as images and videos has recently become an important research subject in computer vision. This paper presents our computation model for predicting image memorability, which is based on a deep learning architecture designed for a classification task. We exploit the…

Cited by 0SourceScholar
2017

Informed source separation via compressive graph signal sampling

ICASSP 2017accepted

We propose a novel informed source separation method for audio object coding based on a recent sampling theory for smooth signals on graphs. Assuming that only one source is active at each time-frequency point, we compute an ideal map indicating which source is active at each time-frequency point at…

Cited by 0SourceScholar
2017

Motion informed audio source separation

ICASSP 2017accepted

In this paper we tackle the problem of single channel audio source separation driven by descriptors of the sounding object's motion. As opposed to previous approaches, motion is included as a soft-coupling constraint within the nonnegative matrix factorization framework. The proposed method is appli…

Cited by 0SourceScholar
2015

Relative group sparsity for non-negative matrix factorization with application to on-the-fly audio source separation

ICASSP 2015accepted

We consider dictionary-based signal decompositions with group sparsity, a variant of structured sparsity. We point out that the group sparsity-inducing constraint alone may not be sufficient in some cases when we know that some bigger groups or so-called supergroups cannot vanish completely. To deal…

Cited by 0SourceScholar