← Search

Han-Gyu Kim

5 accepted papers

2025

MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language Models

ICCV 2025poster

Vision-and-Language Models (VLMs) have shown impressive capabilities on single-turn benchmarks, yet real-world applications often demand more intricate multi-turn dialogues. Existing multi-turn datasets (e.g, MMDU, ConvBench) only partially capture the breadth and depth of conversational scenarios e…

Cited by 0SourcePDFScholar
2024

DialogCC: An Automated Pipeline for Creating High-Quality Multi-Modal Dialogue Dataset

NAACL 2024long

As sharing images in an instant message is a crucial factor, there has been active research on learning an image-text multi-modal dialogue models.However, training a well-generalized multi-modal dialogue model remains challenging due to the low quality and limited diversity of images per dialogue in…

2023

Road Anomaly Segmentation Based on Pixel-wise Logit Variance with Iterative Background Highlighting

ICRA 2023poster

Anomaly segmentation on the urban landscape scene is an important task in autonomous driving. This process exploits a pre-trained semantic segmentation network to estimate anomalous regions. Anomaly segmentation approaches implemented with extra requirements such as out-of-domain data, extra network…

Cited by 0SourcecodeScholar
2021

Proxy Synthesis: Learning with Synthetic Classes for Deep Metric Learning

AAAI 2021technical

One of the main purposes of deep metric learning is to construct an embedding space that has well-generalized embeddings on both seen (training) classes and unseen (test) classes. Most existing works have tried to achieve this using different types of metric objectives and hard sample mining strateg…