← Search

Hyemin Jeong

3 accepted papers

2026

A More Word-like Image Tokenization for MLLMs

CVPR 2026

Modern multimodal large language models (MLLMs) typically keep the language model fixed and train a visual projector that maps the pixels into a sequence of tokens in its embedding space, so that images can be presented in essentially the same form as text. However, the language model has been optim

Cited by 0SourcecodeScholar
2026

TripleSumm: Adaptive Triple-Modality Fusion for Video Summarization

ICLR 2026poster

The exponential growth of video content highlights the importance of video summarization, a task that efficiently extracts key information from long videos. However, existing video summarization studies face inherent limitations in understanding complex, multimodal videos. This limitation stems from…

Cited by 0SourcecodeScholar
2025

K/DA: Automated Data Generation Pipeline for Detoxifying Implicitly Offensive Language in Korean

ACL 2025long

Language detoxification involves removing toxicity from offensive language. While a neutral-toxic paired dataset provides a straightforward approach for training detoxification models, creating such datasets presents several challenges: i) the need for human annotation to build paired data, and ii)…