← Search

Ig-Jae Kim

13 accepted papers

2025

Channel-wise Noise Scheduled Diffusion for Inverse Rendering in Indoor Scenes

CVPR 2025poster

We propose a diffusion-based inverse rendering framework that decomposes a single RGB image into geometry, material, and lighting. Inverse rendering is inherently ill-posed, making it difficult to predict a single accurate solution. To address this challenge, recent generative model-based methods ai…

Cited by 0SourcePDFScholar
2025

Effective SAM Combination for Open-Vocabulary Semantic Segmentation

CVPR 2025poster

Open-vocabulary semantic segmentation aims to assign pixel-level labels to images across an unlimited range of classes. Traditional methods address this by sequentially connecting a powerful mask proposal generator, such as the Segment Anything Model (SAM), with a pre-trained vision-language model l…

Cited by 0SourcePDFScholar
2025

Navigating Label Ambiguity for Facial Expression Recognition in the Wild

AAAI 2025technical

Facial expression recognition (FER) remains a challenging task due to label ambiguity caused by the subjective nature of facial expressions and noisy samples. Additionally, class imbalance, which is common in real-world datasets, further complicates FER. Although many studies have shown impressive i…

2025

VIGFace: Virtual Identity Generation for Privacy-Free Face Recognition Dataset

ICCV 2025poster

Deep learning-based face recognition continues to face challenges due to its reliance on huge datasets obtained from web crawling, which can be costly to gather and raise significant real-world privacy concerns. To address this issue, we propose VIGFace, a novel framework capable of generating synth…

2024

Dual Prototype Attention for Unsupervised Video Object Segmentation

CVPR 2024poster

Unsupervised video object segmentation (VOS) aims to detect and segment the most salient object in videos. The primary techniques used in unsupervised VOS are 1) the collaboration of appearance and motion information; and 2) temporal fusion between different frames. This paper proposes two novel pro…

2024

Few-Shot Neural Radiance Fields under Unconstrained Illumination

AAAI 2024technical

In this paper, we introduce a new challenge for synthesizing novel view images in practical environments with limited input multi-view images and varying lighting conditions. Neural radiance fields (NeRF), one of the pioneering works for this task, demand an extensive set of multi-view images taken…

Cited by 2SourcePDFScholar
2023

MAIR: Multi-View Attention Inverse Rendering With 3D Spatially-Varying Lighting Estimation

CVPR 2023poster

We propose a scene-level inverse rendering framework that uses multi-view images to decompose the scene into geometry, a SVBRDF, and 3D spatially-varying lighting. Because multi-view images provide a variety of information about the scene, multi-view images in object-level inverse rendering have bee…

Cited by 8SourcePDFScholar
2021

Cross-Domain Grouping and Alignment for Domain Adaptive Semantic Segmentation

AAAI 2021technical

Existing techniques to adapt semantic segmentation networks across source and target domains within deep convolutional neural networks (CNNs) deal with all the samples from the two domains in a global or category-aware manner. They do not consider an inter-class variation within the target domain it…

2021

Learning Canonical 3D Object Representation for Fine-Grained Recognition

ICCV 2021poster

We propose a novel framework for fine-grained object recognition that learns to recover object variation in 3D space from a single image, trained on an image collection without using any ground-truth 3D annotation. We accomplish this by representing an object as a composition of 3D shape and its app…

Cited by 16PDFScholar
2020

Cylindrical Convolutional Networks for Joint Object Detection and Viewpoint Estimation

CVPR 2020poster

Existing techniques to encode spatial invariance within deep convolutional neural networks only model 2D transformation fields. This does not account for the fact that objects in a 2D space are a projection of 3D ones, and thus they have limited ability to severe object viewpoint changes. To overcom…

Cited by 19PDFScholar