← Search

Hogun Park

11 accepted papers

2026

Judo: A Juxtaposed Domain-oriented Multimodal Reasoner for Industrial Anomaly QA

ICLR 2026poster

Industrial anomaly detection has been significantly advanced by large multimodal models (LMMs), enabling diverse human instructions beyond detection, particularly through visual-grounded reasoning for better image understanding. However, the lack of domain-specific knowledge of LMMs limits the accur…

Cited by 0SourcecodeScholar
2026

M^3KG-RAG: Multi-hop Multimodal Knowledge Graph-enhanced Retrieval-Augmented Generation

CVPR 2026

Retrieval-Augmented Generation (RAG) has recently been extended to multimodal settings, connecting multimodal large language models (MLLMs) with vast corpora of external knowledge such as multimodal knowledge graphs (MMKGs). Despite their recent success, multimodal RAG in the audio-visual domain rem

Cited by 0SourceScholar
2026

SToRM: Supervised Token Reduction for Multi-Modal LLMs Toward Efficient End-To-End Autonomous Driving

ICRA 2026poster

In autonomous driving, end-to-end(E2E) driving systems that predict control commands directly from sensor data achieved significant advancements. For safe autonomous driving in unexpected scenarios, one may additionally rely on human interventions such as natural language instructions.Using a multi-…

2025

AudioGenX: Explainability on Text-to-Audio Generative Models

AAAI 2025technical

Text-to-audio generation models (TAG) have achieved significant advances in generating audio conditioned on text descriptions. However, a critical challenge lies in the lack of transparency regarding how each textual input impacts the generated audio. To address this issue, we introduce AudioGenX, a…

2025

Enhancing Complex Reasoning in Knowledge Graph Question Answering through Query Graph Approximation

ACL 2025finding

Knowledge-grounded Question Answering (QA) aims to provide answers to structured queries or natural language questions by leveraging Knowledge Graphs (KGs). Existing approaches are mainly divided into Knowledge Graph Question Answering (KGQA) and Complex Query Answering (CQA). Both approaches have l…

Cited by 0SourcePDFScholar
2025

Large Language Models Are Better Logical Fallacy Reasoners with Counterargument, Explanation, and Goal-Aware Prompt Formulation

NAACL 2025findings

The advancement of Large Language Models (LLMs) has greatly improved our ability to process complex language. However, accurately detecting logical fallacies remains a significant challenge. This study presents a novel and effective prompt formulation approach for logical fallacy detection, applicab…

2025

MAMS: Model-Agnostic Module Selection Framework for Video Captioning

AAAI 2025technical

Multi-modal transformers are rapidly gaining attention in video captioning tasks. Existing multi-modal video captioning methods extract a fixed number of frames, but this has critical challenges. If a limited number of frames are extracted, important frames with essential information for caption gen…

2024

Improving Multi-hop Logical Reasoning in Knowledge Graphs with Context-Aware Query Representation Learning

ACL 2024findings

Multi-hop logical reasoning on knowledge graphs is a pivotal task in natural language processing, with numerous approaches aiming to answer First-Order Logic (FOL) queries. Recent geometry (e.g., box, cone) and probability (e.g., beta distribution)-based methodologies have effectively addressed comp…

2024

UNR-Explainer: Counterfactual Explanations for Unsupervised Node Representation Learning Models

ICLR 2024poster

Node representation learning, such as Graph Neural Networks (GNNs), has become one of the important learning methods in machine learning, and the demand for reliable explanation generation is growing. Despite extensive research on explanation generation for supervised node representation learning, e…

Cited by 3SourcePDFScholar