← Search

Byungsoo Ko

10 accepted papers

2025

MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language Models

ICCV 2025poster

Vision-and-Language Models (VLMs) have shown impressive capabilities on single-turn benchmarks, yet real-world applications often demand more intricate multi-turn dialogues. Existing multi-turn datasets (e.g, MMDU, ConvBench) only partially capture the breadth and depth of conversational scenarios e…

Cited by 0SourcePDFScholar
2024

DialogCC: An Automated Pipeline for Creating High-Quality Multi-Modal Dialogue Dataset

NAACL 2024long

As sharing images in an instant message is a crucial factor, there has been active research on learning an image-text multi-modal dialogue models.However, training a well-generalized multi-modal dialogue model remains challenging due to the low quality and limited diversity of images per dialogue in…

2024

Stark: Social Long-Term Multi-Modal Conversation with Persona Commonsense Knowledge

EMNLP 2024finding

Humans share a wide variety of images related to their personal experiences within conversations via instant messaging tools. However, existing works focus on (1) image-sharing behavior in singular sessions, leading to limited long-term social interaction, and (2) a lack of personalized image-sharin…

2024

VVS: Video-to-Video Retrieval with Irrelevant Frame Suppression

AAAI 2024technical

In content-based video retrieval (CBVR), dealing with large-scale collections, efficiency is as important as accuracy; thus, several video-level feature-based studies have actively been conducted. Nevertheless, owing to the severe difficulty of embedding a lengthy and untrimmed video into a single f…

2022

Deep Hash Distillation for Image Retrieval

ECCV 2022poster

"In hash-based image retrieval systems, degraded or transformed inputs usually generate different codes from the original, deteriorating the retrieval accuracy. To mitigate this issue, data augmentation can be applied during training. However, even if augmented samples of an image are similar in rea…

2022

Granularity-Aware Adaptation for Image Retrieval over Multiple Tasks

ECCV 2022poster

"Strong image search models can be learned for a specific domain, ie. set of labels, provided that some labeled images of that domain are available. A practical visual search model, however, should be versatile enough to solve multiple retrieval tasks simultaneously, even if those cover very differe…

Cited by 9SourcePDFScholar
2022

Towards Light-Weight and Real-Time Line Segment Detection

AAAI 2022technical

Previous deep learning-based line segment detection (LSD) suffers from the immense model size and high computational cost for line prediction. This constrains them from real-time inference on computationally restricted environments. In this paper, we propose a real-time and light-weight line segment…

2021

Proxy Synthesis: Learning with Synthetic Classes for Deep Metric Learning

AAAI 2021technical

One of the main purposes of deep metric learning is to construct an embedding space that has well-generalized embeddings on both seen (training) classes and unseen (test) classes. Most existing works have tried to achieve this using different types of metric objectives and hard sample mining strateg…