← Search

Wooyoung Kang

3 accepted papers

2024

Honeybee: Locality-enhanced Projector for Multimodal LLM

CVPR 2024highlight

In Multimodal Large Language Models (MLLMs) a visual projector plays a crucial role in bridging pre-trained vision encoders with LLMs enabling profound visual understanding while harnessing the LLMs' robust capabilities. Despite the importance of the visual projector it has been relatively less expl…

2024

Towards a Complete Benchmark on Video Moment Localization

AISTATS 2024poster

In this paper, we propose and conduct a comprehensive benchmark on moment localization task, which aims to retrieve a segment that corresponds to a text query from a single untrimmed video. Our study starts from an observation that most moment localization papers report experimental results only on…

2023

Noise-Aware Learning from Web-Crawled Image-Text Data for Image Captioning

ICCV 2023poster

Image captioning is one of the straightforward tasks that can take advantage of large-scale web-crawled data which provides rich knowledge about the visual world for a captioning model. However, since web-crawled data contains image-text pairs that are aligned at different levels, the inherent noise…

Cited by 24PDFcodeScholar