← Search

Van-Quang Nguyen

3 accepted papers

2022

GRIT: Faster and Better Image Captioning Transformer Using Dual Visual Features

ECCV 2022poster

"Current state-of-the-art methods for image captioning employ region-based features, as they provide object-level information that is essential to describe the content of images; they are usually extracted by an object detector such as Faster R-CNN. However, they have several issues, such as lack of…

2021

Look Wide and Interpret Twice: Improving Performance on Interactive Instruction-following Tasks

IJCAI 2021poster

There is a growing interest in the community in making an embodied AI agent perform a complicated task while interacting with an environment following natural language directives. Recent studies have tackled the problem using ALFRED, a well-designed dataset for the task, but achieved only very low a…

Cited by 37SourcePDFScholar
2020

Efficient Attention Mechanism for Visual Dialog that can Handle All the Interactions between Multiple Inputs

ECCV 2020poster

It has been a primary concern in recent studies of vision and language tasks to design an effective attention mechanism dealing with interactions between the two modalities. The Transformer has recently been extended and applied to several bi-modal tasks, yielding promising results. For visual dialo…