← Search

Eunji Kim

14 accepted papers

2025

DefectFill: Realistic Defect Generation with Inpainting Diffusion Model for Visual Inspection

CVPR 2025highlight

Developing effective visual inspection models remains challenging due to the scarcity of defect data. While image generation models have been used to synthesize defect images, producing highly realistic defects remains difficult. We propose DefectFill, a novel method for realistic defect generation…

Cited by 0SourcePDFScholar
2025

Interpretable Next-token Prediction via the Generalized Induction Head

NeurIPS 2025poster

While large transformer models excel in predictive performance, their lack of interpretability restricts their usefulness in high-stakes domains. To remedy this, we propose the Generalized Induction-Head Model (GIM), an interpretable model for next-token prediction inspired by the observation of “in…

Cited by 0SourcecodeScholar
2025

Rethinking Training for De-biasing Text-to-Image Generation: Unlocking the Potential of Stable Diffusion

CVPR 2025poster

Recent advancements in text-to-image models, such as Stable Diffusion, show significant demographic biases. Existing de-biasing techniques rely heavily on additional training, which imposes high computational costs and risks of compromising core image generation functionality. This hinders them from…

Cited by 3SourcePDFScholar
2025

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models

ICML 2025poster

Detailed image captioning is essential for tasks like data generation and aiding visually impaired individuals. High-quality captions require a balance between precision and recall, which remains challenging for current multimodal large language models (MLLMs). In this work, we hypothesize that this…

Cited by 0SourcePDFScholar
2024

HandDAGT: A Denoising Adaptive Graph Transformer for 3D Hand Pose Estimation

ECCV 2024poster

"The extraction of keypoint positions from input hand frames, known as 3D hand pose estimation, is crucial for various human-computer interaction applications. However, current approaches often struggle with the dynamic nature of self-occlusion of hands and intra-occlusion with interacting objects.…

2024

Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP

EMNLP 2024finding

A text encoder within Vision-Language Models (VLMs) like CLIP plays a crucial role in translating textual input into an embedding space shared with images, thereby facilitating the interpretative analysis of vision tasks through natural language. Despite the varying significance of different textual…

Cited by 0SourcePDFScholar
2023

Improving Visual Prompt Tuning for Self-supervised Vision Transformers

ICML 2023poster

Visual Prompt Tuning (VPT) is an effective tuning method for adapting pretrained Vision Transformers (ViTs) to downstream tasks. It leverages extra learnable tokens, known as prompts, which steer the frozen pretrained ViTs. Although VPT has demonstrated its applicability with supervised vision trans…

2022

Bridging the Gap Between Classification and Localization for Weakly Supervised Object Localization

CVPR 2022poster

Weakly supervised object localization aims to find a target object region in a given image with only weak supervision, such as image-level labels. Most existing methods use a class activation map (CAM) to generate a localization map; however, a CAM identifies only the most discriminative parts of a…

Cited by 56PDFcodeScholar
2022

Weakly Supervised Semantic Segmentation Using Out-of-Distribution Data

CVPR 2022poster

Weakly supervised semantic segmentation (WSSS) methods are often built on pixel-level localization maps obtained from a classifier. However, training on class labels only, classifiers suffer from the spurious correlation between foreground and background cues (e.g. train and rail), fundamentally bou…

Cited by 127PDFcodeScholar
2021

Anti-Adversarially Manipulated Attributions for Weakly and Semi-Supervised Semantic Segmentation

CVPR 2021poster

Weakly supervised semantic segmentation produces a pixel-level localization from class labels; but a classifier trained on such labels is likely to restrict its focus to a small discriminative region of the target object. AdvCAM is an attribution map of an image that is manipulated to increase the c…

Cited by 307PDFcodeScholar
2021

XProtoNet: Diagnosis in Chest Radiography With Global and Local Explanations

CVPR 2021poster

Automated diagnosis using deep neural networks in chest radiography can help radiologists detect life-threatening diseases. However, existing methods only provide predictions without accurate explanations, undermining the trustworthiness of the diagnostic methods. Here, we present XProtoNet, a globa…

Cited by 144PDFScholar
2019

FickleNet: Weakly and Semi-Supervised Semantic Image Segmentation Using Stochastic Inference

CVPR 2019poster

The main obstacle to weakly supervised semantic image segmentation is the difficulty of obtaining pixel-level information from coarse image-level annotations. Most methods based on image-level annotations use localization maps obtained from the classifier, but these only focus on the small discrimin…

Cited by 557PDFScholar
2019

Frame-to-Frame Aggregation of Active Regions in Web Videos for Weakly Supervised Semantic Segmentation

ICCV 2019poster

When a deep neural network is trained on data with only image-level labeling, the regions activated in each image tend to identify only a small region of the target object. We propose a method of using videos automatically harvested from the web to identify a larger region of the target object by us…

Cited by 45PDFScholar