← Search

Hyunjong Ok

5 accepted papers

2026

AuditoryBench++: Can Language Models Understand Auditory Knowledge without Hearing?

ICASSP 2026oral

Even without directly hearing sounds, humans can effortlessly reason about auditory properties, such as pitch, loudness, or sound-source associations, drawing on auditory commonsense. In contrast, language models often lack this capability, limiting their effectiveness in multimodal interactions. As…

Cited by 0SourcePDFScholar
2025

Imagine to Hear: Auditory Knowledge Generation can be an Effective Assistant for Language Models

ACL 2025finding

Language models pretrained on text-only corpora often struggle with tasks that require auditory commonsense knowledge.Previous work addresses this problem by augmenting the language model to retrieve knowledge from external audio databases.This approach has several limitations, such as the potential…

Cited by 0SourcePDFScholar
2024

SCANNER: Knowledge-Enhanced Approach for Robust Multi-modal Named Entity Recognition of Unseen Entities

NAACL 2024long

Recent advances in named entity recognition (NER) have pushed the boundary of the task to incorporate visual signals, leading to many variants, including multi-modal NER (MNER) or grounded MNER (GMNER). A key challenge to these tasks is that the model should be able to generalize to the entities uns…

Cited by 3SourcePDFScholar
2023

Post-Trained Language Model Adaptive to Extractive Summarization of Long Spoken Documents

ICASSP 2023accepted

The General Meeting Understanding and Generation Challenge challenge track 2 in ICASSP 2023 Signal Processing Grand Challenge aims at extractive summarization of meeting transcripts. The main characteristic of this challenge is abridged as long spoken documents that are extremely difficult to manipu…

Cited by 0SourceScholar