← Search

Yeong Hyeon Gu

4 accepted papers

2026

DTP: Delta-Guided Two Stage Pruning for Mamba-based Multimodal Large Language Models

ICLR 2026poster

Multimodal large language models built on the Mamba architecture offer efficiency advantages, yet remain hampered by redundant visual tokens that inflate inference cost, with the prefill stage accounting for the majority of total inference time. We introduce Delta-guided Two stage Pruning (DTP), a m…

Cited by 0SourceScholar
2024

On Train-Test Class Overlap and Detection for Image Retrieval

CVPR 2024poster

How important is it for training and evaluation sets to not have class overlap in image retrieval? We revisit Google Landmarks v2 clean the most popular training set by identifying and removing class overlap with Revisited Oxford and Paris the most popular training set. By comparing the original and…

Cited by 7SourcePDFScholar
2024

SyncMask: Synchronized Attentional Masking for Fashion-centric Vision-Language Pretraining

CVPR 2024poster

Vision-language models (VLMs) have made significant strides in cross-modal understanding through large-scale paired datasets. However in fashion domain datasets often exhibit a disparity between the information conveyed in image and text. This issue stems from datasets containing multiple images of…

Cited by 9SourcePDFScholar
2023

Conditional Cross Attention Network for Multi-Space Embedding without Entanglement in Only a SINGLE Network

ICCV 2023poster

Many studies in vision tasks have aimed to create effective embedding spaces for single-label object prediction within an image. However, in reality, most objects possess multiple specific attributes, such as shape, color, and length, with each attribute composed of various classes. To apply models…

Cited by 2PDFScholar