← Search

Dong-Jin Kim

14 accepted papers

2026

Adaptive Auxiliary Prompt Blending for Target-Faithful Diffusion Generation

CVPR 2026

Diffusion-based text-to-image (T2I) models have made remarkable progress in generating photorealistic and semantically rich images. However, when the target concepts lie in low-density regions of the training distribution, these models often produce semantically misaligned or structurally inconsiste

Cited by 0SourceScholar
2026

Follow the Saliency: Supervised Saliency for Retrieval-augmented Dense Video Captioning

CVPR 2026

Existing retrieval-augmented approaches for Dense Video Captioning (DVC) often fail to achieve accurate temporal segmentation aligned with true event boundaries, as they rely on heuristic strategies that overlook ground truth event boundaries.The proposed framework, STaRC, overcomes this limitation

Cited by 0SourcecodeScholar
2026

SAIL: Similarity-Aware Guidance and Inter-Caption Augmentation-based Learning for Weakly-Supervised Dense Video Captioning

CVPR 2026

Weakly-Supervised Dense Video Captioning aims to localize and describe events in videos trained only on caption annotations, without temporal boundaries. Prior work introduced an implicit supervision paradigm based on Gaussian masking and complementary captioning. However, existing method focus mere

Cited by 0SourceScholar
2025

Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning

EMNLP 2025

Dense video captioning aims to temporally localize events in video and generate captions for each event. While recent works propose end-to-end models, they suffer from two limitations: (1) applying timestamp supervision only to text while treating all video frames equally, and (2) retrieving caption

2025

ScaleDiff: Higher-Resolution Image Synthesis via Efficient and Model-Agnostic Diffusion

NeurIPS 2025poster

Text-to-image diffusion models often exhibit degraded performance when generating images beyond their training resolution. Recent training-free methods can mitigate this limitation, but they often require substantial computation or are incompatible with recent Diffusion Transformer models. In this p…

Cited by 0SourceScholar
2025

VerbDiff: Text-Only Diffusion Models with Enhanced Interaction Awareness

CVPR 2025poster

Recent large-scale text-to-image diffusion models generate photorealistic images but often struggle to accurately depict interactions between humans and objects due to their limited ability to differentiate various interaction words.In this work, we propose VerbDiff to address the challenge of captu…

2025

ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning

AAAI 2025technical

Recent lightweight image captioning models using retrieved data mainly focus on text prompts. However, previous works only utilize the retrieved text as text prompts, and the visual information relies only on the CLIP visual embedding. Because of this issue, there is a limitation that the image desc…

2024

IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning

EMNLP 2024main

Recent advancements in image captioning have explored text-only training methods to overcome the limitations of paired image-text data. However, existing text-only training methods often overlook the modality gap between using text data during training and employing images during inference. To addre…

2024

Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality

EMNLP 2024main

In this paper, we propose a new method to enhance compositional understanding in pre-trained vision and language models (VLMs) without sacrificing performance in zero-shot multi-modal tasks. Traditional fine-tuning approaches often improve compositional reasoning at the cost of degrading multi-modal…

2023

Generative Bias for Robust Visual Question Answering

CVPR 2023poster

The task of Visual Question Answering (VQA) is known to be plagued by the issue of VQA models exploiting biases within the dataset to make its final prediction. Various previous ensemble based debiasing methods have been proposed where an additional model is purposefully trained to be biased in orde…

2023

Self-Sufficient Framework for Continuous Sign Language Recognition

ICASSP 2023accepted

The goal of this work is to develop self-sufficient framework for Continuous Sign Language Recognition (CSLR) that addresses key issues of sign language recognition. These include the need for complex multi-scale features such as hands, face, and mouth for understanding, and absence of frame-level a…

Cited by 0SourceScholar
2022

DASO: Distribution-Aware Semantics-Oriented Pseudo-Label for Imbalanced Semi-Supervised Learning

CVPR 2022poster

The capability of the traditional semi-supervised learning (SSL) methods is far from real-world application due to severely biased pseudo-labels caused by (1) class imbalance and (2) class distribution mismatch between labeled and unlabeled data. This paper addresses such a relatively under-explored…

Cited by 119PDFcodeScholar
2021

LabOR: Labeling Only if Required for Domain Adaptive Semantic Segmentation

ICCV 2021poster

Unsupervised Domain Adaptation (UDA) for semantic segmentation has been actively studied to mitigate the domain gap between label-rich source data and unlabeled target data. Despite these efforts, UDA still has a long way to go to reach the fully supervised performance. To this end, we propose a Lab…

Cited by 55PDFScholar
2019

Dense Relational Captioning: Triple-Stream Networks for Relationship-Based Captioning

CVPR 2019poster

Our goal in this work is to train an image captioning model that generates more dense and informative captions. We introduce "relational captioning," a novel image captioning task which aims to generate multiple captions with respect to relational information between objects in an image. Relational…

Cited by 112PDFcodeScholar