← Search

Hyungwook Choi

3 accepted papers

2026

A More Word-like Image Tokenization for MLLMs

CVPR 2026

Modern multimodal large language models (MLLMs) typically keep the language model fixed and train a visual projector that maps the pixels into a sequence of tokens in its embedding space, so that images can be presented in essentially the same form as text. However, the language model has been optim

Cited by 0SourcecodeScholar
2026

Fine-Grained Multi Image Object Hallucination Benchmark

CVPR 2026

Multimodal Large Language Models (MLLMs) are increasingly deployed in multi-image scenarios requiring complex reasoning across visual contexts. However, current MLLMs remain fundamentally limited by object hallucination--generating plausible yet factually inconsistent descriptions about objects. Exi

Cited by 0SourceScholar