← Search

Jaeyoung Lee

9 accepted papers

2026

Do Language Models Associate Sound with Meaning? A Multimodal Study of Sound Symbolism

AAAI 2026technical

Sound symbolism is a linguistic concept that refers to non-arbitrary associations between phonetic forms and their meanings. We suggest that this can be a compelling probe into how Multimodal Large Language Models (MLLMs) interpret auditory information in human languages. We investigate MLLMs

Cited by 0SourcePDFScholar
2025

AI Debate Aids Assessment of Controversial Claims

NeurIPS 2025poster

As AI grows more powerful, it will increasingly shape how we understand the world. But with this influence comes the risk of amplifying misinformation and deepening social divides—especially on consequential topics where factual accuracy directly impacts well-being. Scalable Oversight aims to ensure…

Cited by 0SourceScholar
2025

From Curiosity to Clarity : Exploring the Impact of Consecutive Why-Questions

NAACL 2025findings

Humans attempt to understand the real world by asking the fundamental question ”Why?” when faced with incomprehensible situations in everyday life. Such why-questions provide essential knowledge that can help in understanding these situations. In this study, we conducted an end-to-end process to ver…

2025

Leveraging IPA and Articulatory Features as Effective Inductive Biases for Multilingual ASR Training

ICASSP 2025accepted

In recent advancements in end-to-end ASR, large-scale self-supervised or weakly supervised models have achieved a significant milestone. However, it remains challenging to train consistently high-performing multilingual models, transferable to languages without much resource. In this study, we propo…

Cited by 0SourceScholar
2025

SynAD: Enhancing Real-World End-to-End Autonomous Driving Models through Synthetic Data Integration

ICCV 2025poster

Recent advancements in deep learning and the availability of high-quality real-world driving datasets have propelled end-to-end autonomous driving. Despite this progress, relying solely on real-world data limits the variety of driving scenarios for training. Synthetic scenario generation has emerged…

Cited by 0SourcePDFScholar
2024

ESG-Kor: A Korean Dataset for ESG-related Information Extraction and Practical Use Cases

EMNLP 2024finding

With the expansion of pre-trained language model usage in recent years, the importance of datasets for performing tasks in specialized domains has significantly increased. Therefore, we have built a Korean dataset called ESG-Kor to automatically extract Environmental, Social, and Governance (ESG) in…

2024

Hierarchical Graph Convolutional Network Approach for Detecting Low-Quality Documents

COLING 2024main

Consistency within a document is a crucial feature indicative of its quality. Recently, within the vast amount of information produced across various media, there exists a significant number of low-quality documents that either lack internal consistency or contain content utterly unrelated to their…

2024

How to Train Your Fact Verifier: Knowledge Transfer with Multimodal Open Models

EMNLP 2024finding

Given the growing influx of misinformation across news and social media, there is a critical need for systems that can provide effective real-time verification of news claims. Large language or multimodal model based verification has been proposed to scale up online policing mechanisms for mitigatin…

Cited by 0SourcePDFScholar