← Search

Kwanyong Park

12 accepted papers

2025

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning

ICCV 2025poster

An image captioning model flexibly switching its language pattern, e.g., descriptiveness and length, should be useful since it can be applied to diverse applications. However, despite the dramatic improvement in generative vision-language models, fine-grained control over the properties of generated…

2024

KOALA: Empirical Lessons Toward Memory-Efficient and Fast Diffusion Models for Text-to-Image Synthesis

NeurIPS 2024poster

As text-to-image (T2I) synthesis models increase in size, they demand higher inference costs due to the need for more expensive GPUs with larger memory, which makes it challenging to reproduce these models in addition to the restricted access to training datasets. Our study aims to reduce these infe…

Cited by 2SourcePDFScholar
2024

MTMMC: A Large-Scale Real-World Multi-Modal Camera Tracking Benchmark

CVPR 2024poster

Multi-target multi-camera tracking is a crucial task that involves identifying and tracking individuals over time using video streams from multiple cameras. This task has practical applications in various fields such as visual surveillance crowd behavior analysis and anomaly detection. However due t…

Cited by 1SourcePDFScholar
2024

Weak-to-Strong Compositional Learning from Generative Models for Language-based Object Detection

ECCV 2024poster

"Vision-language (VL) models often exhibit a limited understanding of complex expressions of visual objects (, attributes, shapes, and their relations), given complex and diverse language queries. Traditional approaches attempt to improve VL models using hard negative synthetic text, but their effec…

Cited by 1SourcePDFScholar
2023

Bidirectional Domain Mixup for Domain Adaptive Semantic Segmentation

AAAI 2023technical

Mixup provides interpolated training samples and allows the model to obtain smoother decision boundaries for better generalization. The idea can be naturally applied to the domain adaptation task, where we can mix the source and target samples to obtain domain-mixed samples for better adaptation. Ho…

2023

Test-Time Adaptation in the Dynamic World With Compound Domain Knowledge Management

RA-L 2023

Prior to the deployment of robotic systems, pre-training the deep-recognition models on all potential visual cases is infeasible in practice. Hence, test-time adaptation (TTA) allows the model to adapt itself to novel environments and improve its performance during test time (i.e., lifelong adaptati

Cited by 9SourceScholar
2022

Bridging Images and Videos: A Simple Learning Framework for Large Vocabulary Video Object Detection

ECCV 2022poster

"Scaling object taxonomies is one of the important steps toward a robust real-world deployment of recognition systems. We have faced remarkable progress in images since the introduction of the LVIS benchmark. To continue this success in videos, a new video benchmark, TAO, was recently presented. Giv…

Cited by 8SourcePDFScholar
2021

LabOR: Labeling Only if Required for Domain Adaptive Semantic Segmentation

ICCV 2021poster

Unsupervised Domain Adaptation (UDA) for semantic segmentation has been actively studied to mitigate the domain gap between label-rich source data and unlabeled target data. Despite these efforts, UDA still has a long way to go to reach the fully supervised performance. To this end, we propose a Lab…

Cited by 55PDFScholar
2020

Discover, Hallucinate, and Adapt: Open Compound Domain Adaptation for Semantic Segmentation

NeurIPS 2020poster

Unsupervised domain adaptation (UDA) for semantic segmentation has been attracting attention recently, as it could be beneficial for various label-scarce real-world scenarios (e.g., robot control, autonomous driving, medical imaging, etc.). Despite the significant progress in this field, current wor…

Cited by 39SourcePDFScholar