← Search

Yu-Jung Heo

8 accepted papers

2025

Evaluating Visual and Cultural Interpretation: The K-Viscuit Benchmark with Human-VLM Collaboration

ACL 2025long

To create culturally inclusive vision-language models (VLMs), developing a benchmark that tests their ability to address culturally relevant questions is essential. Existing approaches typically rely on human annotators, making the process labor-intensive and creating a cognitive burden in generatin…

Cited by 0SourcePDFScholar
2024

BI-MDRG: Bridging Image History in Multimodal Dialogue Response Generation

ECCV 2024poster

"Multimodal Dialogue Response Generation (MDRG) is a recently proposed task where the model needs to generate responses in texts, images, or a blend of both based on the dialogue context. Due to the lack of a large-scale dataset specifically for this task and the benefits of leveraging powerful pre-…

2024

Structure-Aware Multimodal Sequential Learning for Visual Dialog

AAAI 2024technical

With the ability to collect vast amounts of image and natural language data from the web, there has been a remarkable advancement in Large-scale Language Models (LLMs). This progress has led to the emergence of chatbots and dialogue systems capable of fluent conversations with humans. As the variety…

Cited by 1SourcePDFScholar
2024

Translation Deserves Better: Analyzing Translation Artifacts in Cross-lingual Visual Question Answering

ACL 2024findings

Building a reliable visual question answering (VQA) system across different languages is a challenging problem, primarily due to the lack of abundant samples for training. To address this challenge, recent studies have employed machine translation systems for the cross-lingual VQA task. This involve…

2022

Hypergraph Transformer: Weakly-Supervised Multi-hop Reasoning for Knowledge-based Visual Question Answering

ACL 2022long

Knowledge-based visual question answering (QA) aims to answer a question which requires visually-grounded external knowledge beyond image content itself. Answering complex questions that require multi-hop reasoning under weak supervision is considered as a challenging problem since i) no supervision…

2021

DramaQA: Character-Centered Video Story Understanding with Hierarchical QA

AAAI 2021technical

Despite recent progress on computer vision and natural language processing, developing a machine that can understand video story is still hard to achieve due to the intrinsic difficulty of video story. Moreover, researches on how to evaluate the degree of video understanding based on human cognitive…

2020

Hypergraph Attention Networks for Multimodal Learning

CVPR 2020poster

One of the fundamental problems that arise in multimodal learning tasks is the disparity of information levels between different modalities. To resolve this problem, we propose Hypergraph Attention Networks (HANs), which define a common semantic space among the modalities with symbolic graphs and ex…

Cited by 120PDFcodeScholar
2018

Answerer in Questioner's Mind: Information Theoretic Approach to Goal-Oriented Visual Dialog

NeurIPS 2018spotlight

Goal-oriented dialog has been given attention due to its numerous applications in artificial intelligence. Goal-oriented dialogue tasks occur when a questioner asks an action-oriented question and an answerer responds with the intent of letting the questioner know a correct action to take. To ask t…

Cited by 43SourcePDFScholar