← Search

Minuk Ma

4 accepted papers

2025

Learning to See through Sound: From VggCaps to Multi2Cap for Richer Automated Audio Captioning

EMNLP 2025

Automated Audio Captioning (AAC) aims to generate natural language descriptions of audio content, enabling machines to interpret and communicate complex acoustic scenes. However, current AAC datasets often suffer from short and simplistic captions, limiting model expressiveness and semantic depth. T

Cited by 0SourcePDFScholar
2020

Modality Shifting Attention Network for Multi-Modal Video Question Answering

CVPR 2020poster

This paper considers a network referred to as Modality Shifting Attention Network (MSAN) for Multimodal Video Question Answering (MVQA) task. MSAN decomposes the task into two sub-tasks: (1) localization of temporal moment relevant to the question, and (2) accurate prediction of the answer based on…

Cited by 102PDFScholar
2020

VLANet: Video-Language Alignment Network for Weakly-Supervised Video Moment Retrieval

ECCV 2020poster

Video Moment Retrieval (VMR) is a task to localize the temporal moment in untrimmed video specified by natural language query. For VMR, several methods that require full supervision for training have been proposed. Unfortunately, acquiring a large number of training videos with labeled temporal boun…

Cited by 99SourcePDFScholar
2019

Progressive Attention Memory Network for Movie Story Question Answering

CVPR 2019poster

This paper proposes the progressive attention memory network (PAMN) for movie story question answering (QA). Movie story QA is challenging compared to VQA in two aspects: (1) pinpointing the temporal parts relevant to answer the question is difficult as the movies are typically longer than an hour,…

Cited by 96PDFScholar