← Search

Jung Uk Kim

22 accepted papers

2026

Do We Need Perfect Data? Leveraging Noise for Domain Generalized Segmentation

AAAI 2026technical

Domain generalization in semantic segmentation faces challenges from domain shifts, particularly under adverse conditions. While diffusion-based data generation methods show promise, they introduce inherent misalignment between generated images and semantic masks. This paper presents FLEX-Seg (FLexi

Cited by 0SourcePDFScholar
2026

Learning from Oblivion: Predicting Knowledge-Overflowed Weights via Retrodiction of Forgetting

CVPR 2026

Pre-trained weights have become a cornerstone of modern deep learning, enabling efficient knowledge transfer and improving downstream task performance, especially in data-scarce scenarios. However, a fundamental question remains: how can we obtain better pre-trained weights that encapsulate more kno

Cited by 0SourcecodeScholar
2026

Leveraging Textual Compositional Reasoning for Robust Change Captioning

AAAI 2026technical

Change captioning aims to describe changes between a pair of images. However, existing works rely on visual features alone, which often fail to capture subtle but meaningful changes because they lack the ability to represent explicitly structured information such as object relationships and composit

Cited by 0SourcePDFScholar
2026

See, Rank, and Filter: Important Word-Aware Clip Filtering via Scene Understanding for Moment Retrieval and Highlight Detection

AAAI 2026technical

Video moment retrieval (MR) and highlight detection (HD) with natural language queries aim to localize relevant moments and key highlights in a video clips. However, existing methods overlook the importance of individual words, treating the entire text query and video clips as a black-box, which hin

Cited by 0SourcePDFScholar
2026

Task Prototype-Based Knowledge Retrieval for Multi-Task Learning from Partially Annotated Data

AAAI 2026technical

Multi-task learning (MTL) is critical in real-world applications such as autonomous driving and robotics, enabling simultaneous handling of diverse tasks. However, obtaining fully annotated data for all tasks is impractical due to labeling costs. Existing methods for partially labeled MTL typically

Cited by 0SourcePDFScholar
2025

Multispectral Pedestrian Detection with Sparsely Annotated Label

AAAI 2025technical

Although existing Sparsely Annotated Object Detection (SAOD) approches have made progress in handling sparsely annotated environments in multispectral domain, where only some pedestrians are annotated, they still have the following limitations: (i) they lack considerations for improving the quality…

2025

Object-aware Sound Source Localization via Audio-Visual Scene Understanding

CVPR 2025poster

Audio-visual sound source localization task aims to spatially localize sound-making objects within visual scenes by integrating visual and audio cues. However, existing methods struggle with accurately localizing sound-making objects in complex scenes, particularly when visually similar silent objec…

2025

Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight Detection

AAAI 2025technical

The goal of video moment retrieval and highlight detection is to identify specific segments and highlights based on a given text query. With the rapid growth of video content and the overlap between these tasks, recent works have addressed both simultaneously. However, they still struggle to fully c…

2024

Enhancing Audio-Visual Question Answering with Missing Modality via Trans-Modal Associative Learning

ICASSP 2024accepted

We present a novel method for Audio-Visual Question Answering (AVQA) in real-world scenarios where one modality (audio or visual) can be missing. Inspired by human cognitive processes, we introduce a Trans-Modal Associative (TMA) memory that recalls missing modal information (i.e., pseudo modal feat…

Cited by 0SourceScholar
2024

Learning to Visually Localize Sound Sources from Mixtures without Prior Source Knowledge

CVPR 2024poster

The goal of the multi-sound source localization task is to localize sound sources from the mixture individually. While recent multi-sound source localization methods have shown improved performance they face challenges due to their reliance on prior information about the number of objects to be sepa…

2023

Online Class Incremental Learning on Stochastic Blurry Task Boundary via Mask and Visual Prompt Tuning

ICCV 2023poster

Continual learning aims to learn a model from a continuous stream of data, but it mainly assumes a fixed number of data and tasks with clear task boundaries. However, in real-world scenarios, the number of input data and tasks is constantly changing in a statistical way, not a static way. Although r…

Cited by 27PDFcodeScholar
2023

Similarity Relation Preserving Cross-Modal Learning for Multispectral Pedestrian Detection Against Adversarial Attacks

ICASSP 2023accepted

Although multispectral pedestrian detection studies have shown remarkable detection performances, they are still vulnerable to adversarial attacks. We see the similarity relations between object candidates were not maintained because of the adversarial attacks, resulting in performance degradation.…

Cited by 0SourceScholar
2022

Robust Thermal Infrared Pedestrian Detection By Associating Visible Pedestrian Knowledge

ICASSP 2022accepted

Recently, pedestrian detection on thermal infrared images has shown the robust pedestrian detection performance. In this paper, we propose a novel thermal infrared pedestrian detection framework which can associate and utilize the complementary pedestrian knowledge from visible images. Motivated by…

Cited by 0SourceScholar
2022

Towards Versatile Pedestrian Detector with Multisensory-Matching and Multispectral Recalling Memory

AAAI 2022technical

Recently, automated surveillance cameras can change a visible sensor and a thermal sensor for all-day operation. However, existing single-modal pedestrian detectors mainly focus on detecting pedestrians in only one specific modality (i.e., visible or thermal), so they cannot cope with other modal in…

Cited by 30SourcePDFScholar
2021

Towards Robust Training of Multi-Sensor Data Fusion Network Against Adversarial Examples in Semantic Segmentation

ICASSP 2021accepted

The success of multi-sensor data fusions in deep learning appears to be attributed to the use of complementary information among multiple sensor datasets. Compared to their predictive performance, relatively less attention has been devoted to the adversarial robustness of multi-sensor data fusion mo…

Cited by 0SourceScholar
2020

SACA Net: Cybersickness Assessment of Individual Viewers for VR Content via Graph-based Symptom Relation Embedding

ECCV 2020poster

Recently, cybersickness assessment for VR content is required to deal with viewing safety issues. Assessing physical symptoms of individual viewers is challenging but important to provide detailed and personalized guides for viewing safety. In this paper, we propose a novel symptom-aware cybersickne…

Cited by 8SourcePDFScholar
2020

Structure Boundary Preserving Segmentation for Medical Image With Ambiguous Boundary

CVPR 2020poster

In this paper, we propose a novel image segmentation method to tackle two critical problems of medical image, which are (i) ambiguity of structure boundary in the medical image domain and (ii) uncertainty of the segmented region without specialized domain knowledge. To solve those two problems in au…

Cited by 169PDFScholar
2020

Towards High-Performance Object Detection: Task-Specific Design Considering Classification and Localization Separation

ICASSP 2020accepted

Object detection performs two tasks (classification and localization) simultaneously. Two tasks share a similarity: they need robust features that effectively represent the visual appearance of the objects. However, two tasks also have different properties. First, classification mainly requires feat…

Cited by 0SourceScholar