← Search

Soyeon Caren Han

10 accepted papers

2025

3M-Game: Multi-Modal Multi-Task Multi-Teacher Learning for Game Event Detection (Student Abstract)

AAAI 2025technical

Esports has rapidly emerged as a global phenomenon with an ever-expanding audience on livestream platforms. However, due to the complex nature of the game, it becomes challenging for newcomers to comprehend the gaming situation. This research introduces a 3M-Game that integrates multi-modal (MM) inf…

2025

Multimodal Commonsense Knowledge Distillation for Visual Question Answering (Student Abstract)

AAAI 2025technical

Existing Multimodal Large Language Models (MLLMs) and Visual Language Pretrained Models (VLPMs) have shown remarkable performances in general Visual Question Answering (VQA). However, these models struggle with VQA questions that require external commonsense knowledge due to the challenges in genera…

2025

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding

IJCAI 2025

Visually Rich Document Understanding (VRDU) has emerged as a critical field in document intelligence, enabling automated extraction of key information from complex documents across domains such as medical, financial, and educational applications. However, form-like documents pose unique challenges d

Cited by 0SourcePDFScholar
2024

MMVQA: A Comprehensive Dataset for Investigating Multipage Multimodal Information Retrieval in PDF-based Visual Question Answering

IJCAI 2024poster

Document Question Answering (QA) presents a challenge in understanding visually-rich documents (VRD), particularly with lengthy textual content. Existing studies primarily focus on real-world documents with sparse text, while challenges persist in comprehending the hierarchical semantic relations am…

2024

PEACH: Pretrained-Embedding Explanation across Contextual and Hierarchical Structure

IJCAI 2024poster

In this work, we propose a novel tree-based explanation technique, PEACH (Pretrained-embedding Explanation Across Contextual and Hierarchical Structure), that can explain how text-based documents are classified by using any pretrained contextual embeddings in a tree-based human-interpretable manner.…

2023

In-Game Toxic Language Detection: Shared Task and Attention Residuals (Student Abstract)

AAAI 2023technical

In-game toxic language becomes the hot potato in the gaming industry and community. There have been several online game toxicity analysis frameworks and models proposed. However, it is still challenging to detect toxicity due to the nature of in-game chat, which has extremely short length. In this p…

Cited by 1SourcePDFScholar
2022

Doc-GCN: Heterogeneous Graph Convolutional Networks for Document Layout Analysis

COLING 2022main

Recognizing the layout of unstructured digital documents is crucial when parsing the documents into the structured, machine-readable format for downstream applications. Recent studies in Document Layout Analysis usually rely on visual cues to understand documents while ignoring other information, su…

2022

K-MHaS: A Multi-label Hate Speech Detection Dataset in Korean Online News Comment

COLING 2022main

Online hate speech detection has become an important issue due to the growth of online content, but resources in languages other than English are extremely limited. We introduce K-MHaS, a new multi-label dataset for hate speech detection that effectively handles Korean language patterns. The dataset…

2022

Understanding Attention for Vision-and-Language Tasks

COLING 2022main

Attention mechanism has been used as an important component across Vision-and-Language(VL) tasks in order to bridge the semantic gap between visual and textual features. While attention has been widely used in VL tasks, it has not been examined the capability of different attention alignment calcula…