← Search

Seong Tae Kim

15 accepted papers

2026

FedDAP: Domain-Aware Prototype Learning for Federated Learning under Domain Shift

CVPR 2026

Federated Learning (FL) enables decentralized model training across multiple clients without exposing private data, making it ideal for privacy-sensitive applications. However, in real-world FL scenarios, clients often hold data from distinct domains, leading to severe domain shift and degraded glob

Cited by 0SourcecodeScholar
2026

Leveraging Textual Compositional Reasoning for Robust Change Captioning

AAAI 2026technical

Change captioning aims to describe changes between a pair of images. However, existing works rely on visual features alone, which often fail to capture subtle but meaningful changes because they lack the ability to represent explicitly structured information such as object relationships and composit

Cited by 0SourcePDFScholar
2025

Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition

NeurIPS 2025spotlight

Effective explanations of video action recognition models should disentangle how movements unfold over time from the surrounding spatial context. However, existing methods—based on saliency—produce entangled explanations, making it unclear whether predictions rely on motion or spatial context. Langu…

Cited by 0SourceScholar
2025

HiCM²: Hierarchical Compact Memory Modeling for Dense Video Captioning

AAAI 2025technical

With the growing demand for solutions to real-world video challenges, interest in dense video captioning (DVC) has been on the rise. DVC involves the automatic captioning and localization of untrimmed videos. Several studies highlight the challenges of DVC and introduce improved methods utilizing pr…

Cited by 1SourcePDFScholar
2025

LLaVA Needs More Knowledge: Retrieval Augmented Natural Language Generation with Knowledge Graph for Explaining Thoracic Pathologies

AAAI 2025technical

Generating Natural Language Explanations (NLEs) for model predictions on medical images, particularly those depicting thoracic pathologies, remains a critical and challenging task. Existing methodologies often struggle due to general models' insufficient domain-specific medical knowledge and privacy…

2025

When Will It Fail?: Anomaly to Prompt for Forecasting Future Anomalies in Time Series

ICML 2025poster

Recently, forecasting future abnormal events has emerged as an important scenario to tackle realworld necessities. However, the solution of predicting specific future time points when anomalies will occur, known as Anomaly Prediction (AP), remains under-explored. Existing methods dealing with time s…

Cited by 10SourcePDFScholar
2024

Do You Remember? Dense Video Captioning with Cross-Modal Memory Retrieval

CVPR 2024poster

There has been significant attention to the research on dense video captioning which aims to automatically localize and caption all events within untrimmed video. Several studies introduce methods by designing dense video captioning as a multitasking problem of event localization and event captionin…

2024

WWW: A Unified Framework for Explaining What Where and Why of Neural Networks by Interpretation of Neuron Concepts

CVPR 2024poster

Recent advancements in neural networks have showcased their remarkable capabilities across various domains. Despite these successes the "black box" problem still remains. To address this we propose a novel framework WWW that offers the 'what' 'where' and 'why' of the neural network decisions in huma…

2023

LINe: Out-of-Distribution Detection by Leveraging Important Neurons

CVPR 2023poster

It is important to quantify the uncertainty of input samples, especially in mission-critical domains such as autonomous driving and healthcare, where failure predictions on out-of-distribution (OOD) data are likely to cause big problems. OOD detection problem fundamentally begins in that the model c…

2021

Fine-Grained Neural Network Explanation by Identifying Input Features with Predictive Information

NeurIPS 2021poster

One principal approach for illuminating a black-box neural network is feature attribution, i.e. identifying the importance of input features for the network’s prediction. The predictive information of features is recently proposed as a proxy for the measure of their importance. So far, the predictiv…

2021

Neural Response Interpretation Through the Lens of Critical Pathways

CVPR 2021poster

Is critical input information encoded in specific sparse pathways within the neural network? In this work, we discuss the problem of identifying these critical pathways and subsequently leverage them for interpreting the network's response to an input. The pruning objective --- selecting the smalles…

Cited by 41PDFcodeScholar
2020

Force-Ultrasound Fusion: Bringing Spine Robotic-US to the Next "Level"

RA-L 2020

Spine injections are commonly performed in several clinical procedures. The localization of the target vertebral level (i.e. the position of a vertebra in a spine) is typically done by back palpation or under X-ray guidance, yielding either higher chances of procedure failure or exposure to ionizing

Cited by 40SourceScholar
2020

Towards High-Performance Object Detection: Task-Specific Design Considering Classification and Localization Separation

ICASSP 2020accepted

Object detection performs two tasks (classification and localization) simultaneously. Two tasks share a similarity: they need robust features that effectively represent the visual appearance of the objects. However, two tasks also have different properties. First, classification mainly requires feat…

Cited by 0SourceScholar
2018

Facial Dynamics Interpreter Network: What are the Important Relations between Local Dynamics for Facial Trait Estimation?

ECCV 2018poster

Human face analysis is an important task in computer vision. According to cognitive-psychological studies, facial dynamics could provide crucial cues for face analysis. The motion of a facial local region in facial expression is related to the motion of other facial local regions. In this paper, a n…

Cited by 7SourcePDFScholar