← Search

Junho Kim

49 accepted papers

2026

Aligned Novel View Image and Geometry Synthesis via Cross-modal Attention Instillation

ICLR 2026poster

We introduce a diffusion-based framework that generates aligned novel view images and geometries via a warping‐and‐inpainting methodology. Unlike prior methods that require dense posed images or pose-embedded generative models limited to in‐domain views, our method leverages off‐the‐shelf geometry p…

Cited by 0SourcecodeScholar
2026

Generating Humanless Environment Walkthroughs from Egocentric Walking Tour Videos

CVPR 2026

Egocentric walking tour videos provide a rich source of image data to develop rich and diverse visual models of environments around the world. However, the significant presence of humans in frames of these videos due to crowds and eye-level camera perspectives mitigates their usefulness in environme

Cited by 0SourceScholar
2026

Point2Act: Efficient 3D Distillation of Multimodal LLMs for Zero-Shot Context-Aware Grasping

ICRA 2026poster

We propose Point2Act, which directly retrieves the 3D action point relevant to a contextually described task, leveraging Multimodal Large Language Models (MLLMs). Foundation models have opened the possibility for generalist robots that can perform a zero-shot task following natural language descript…

2025

CapeLLM: Support-Free Category-Agnostic Pose Estimation with Multimodal Large Language Models

ICCV 2025poster

Category-agnostic pose estimation (CAPE) has traditionally relied on support images with annotated keypoints, a process that is often cumbersome and may fail to fully capture the necessary correspondences across diverse object categories. Recent efforts have explored the use of text queries, leverag…

2025

Connecting the Knowledge Dots: Retrieval-augmented Knowledge Connection for Commonsense Reasoning

EMNLP 2025

While large language models (LLMs) have achieved remarkable performance across various natural language processing (NLP) tasks, LLMs exhibit a limited understanding of commonsense reasoning due to the necessity of implicit knowledge that is rarely expressed in text. Recently, retrieval-augmented lan

Cited by 0SourcePDFScholar
2025

Enhancing Creative Generation on Stable Diffusion-based Models

CVPR 2025poster

Recent text-to-image generative models, particularly Stable Diffusion and its distilled variants, have achieved impressive fidelity and strong text-image alignment. However, their creative generation capacity remains limited, as simply adding the term "creative" to prompts often fails to yield genui…

2025

Incorporating Domain Knowledge into Materials Tokenization

ACL 2025long

While language models are increasingly utilized in materials science, typical models rely on frequency-centric tokenization methods originally developed for natural language processing. However, these methods frequently produce excessive fragmentation and semantic loss, failing to maintain the struc…

Cited by 0SourcePDFScholar
2025

Learning 3D Scene Analogies with Neural Contextual Scene Maps

ICCV 2025poster

Understanding scene contexts is crucial for machines to perform tasks and adapt prior knowledge in unseen or noisy 3D environments. As data-driven learning is intractable to comprehensively encapsulate diverse ranges of layouts and open spaces, we propose teaching machines to identify relational com…

2025

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis

CVPR 2025poster

Despite advances in Large Multi-modal Models, applying them to long and untrimmed video content remains challenging due to limitations in context length and substantial memory overhead. These constraints often lead to significant information loss and reduced relevance in the model responses. With th…

Cited by 3SourcePDFScholar
2025

StyleKeeper: Prevent Content Leakage using Negative Visual Query Guidance

ICCV 2025poster

In the domain of text-to-image generation, diffusion models have emerged as powerful tools. Recently, studies on visual prompting, where images are used as prompts, have enabled more precise control over style and content. However, existing methods often suffer from content leakage, where undesired…

Cited by 0SourcePDFScholar
2025

“Going to a trap house” conveys more fear than “Going to a mall”: Benchmarking Emotion Context Sensitivity for LLMs

EMNLP 2025

Emotion context sensitivity—the ability to adjust emotional responses based on contexts—is a core component of human emotional intelligence. For example, being told, “You can come with me if you want,” may elicit joy if the destination is a mall, but provoke fear if the destination is a trap house.

Cited by 0SourcePDFScholar
2024

CODE: Contrasting Self-generated Description to Combat Hallucination in Large Multi-modal Models

NeurIPS 2024poster

Large Multi-modal Models (LMMs) have recently demonstrated remarkable abilities in visual context understanding and coherent response generation. However, alongside these advancements, the issue of hallucinations has emerged as a significant challenge, producing erroneous responses that are unrelate…

Cited by 52SourcePDFScholar
2024

Coconut: Contextualized Commonsense Unified Transformers for Graph-Based Commonsense Augmentation of Language Models

ACL 2024findings

In this paper, we introduce COCONUT to effectively guide the contextualization of structured commonsense knowledge based on largelanguage models. COCONUT employs a contextualized knowledge prompting scheme to gather high-quality contextualization examplesfrom a large language model. These examples a…

Cited by 1SourcePDFScholar
2024

Direct Unlearning Optimization for Robust and Safe Text-to-Image Models

NeurIPS 2024poster

Recent advancements in text-to-image (T2I) models have greatly benefited from large-scale datasets, but they also pose significant risks due to the potential generation of unsafe content. To mitigate this issue, researchers proposed unlearning techniques that attempt to induce the model to unlearn p…

Cited by 13SourcePDFScholar
2024

Let 2D Diffusion Model Know 3D-Consistency for Robust Text-to-3D Generation

ICLR 2024poster

Text-to-3D generation has shown rapid progress in recent days with the advent of score distillation sampling (SDS), a methodology of using pretrained text-to-2D diffusion models to optimize a neural radiance field (NeRF) in a zero-shot setting. However, the lack of 3D awareness in the 2D diffusion m…

2024

MELT: Materials-aware Continued Pre-training for Language Model Adaptation to Materials Science

EMNLP 2024finding

We introduce a novel continued pre-training method, MELT (MatEriaLs-aware continued pre-Training), specifically designed to efficiently adapt the pre-trained language models (PLMs) for materials science. Unlike previous adaptation strategies that solely focus on constructing domain-specific corpus,…

Cited by 2SourcePDFScholar
2024

Towards Robust and Generalized Parameter-Efficient Fine-Tuning for Noisy Label Learning

ACL 2024long

Parameter-efficient fine-tuning (PEFT) has enabled the efficient optimization of cumbersome language models in real-world settings. However, as datasets in such environments often contain noisy labels that adversely affect performance, PEFT methods are inevitably exposed to noisy labels. Despite thi…

Cited by 3SourcePDFScholar
2024

What if...?: Thinking Counterfactual Keywords Helps to Mitigate Hallucination in Large Multi-modal Models

EMNLP 2024finding

This paper presents a way of enhancing the reliability of Large Multi-modal Models (LMMs) in addressing hallucination, where the models generate cross-modal inconsistent responses. Without additional training, we propose Counterfactual Inception, a novel method that implants counterfactual thinking…

2023

Calibrating Panoramic Depth Estimation for Practical Localization and Mapping

ICCV 2023poster

The absolute depth values of surrounding environments provide crucial cues for various assistive technologies, such as localization, navigation, and 3D structure estimation. We propose that accurate depth estimated from panoramic images can serve as a powerful and light-weight input for a wide range…

Cited by 2PDFcodeScholar
2023

Client-Customized Adaptation for Parameter-Efficient Federated Learning

ACL 2023findings

Despite the versatility of pre-trained language models (PLMs) across domains, their large memory footprints pose significant challenges in federated learning (FL), where the training model has to be distributed between a server and clients. One potential solution to bypass such constraints might be…

Cited by 22SourcePDFScholar
2023

Demystifying Causal Features on Adversarial Examples and Causal Inoculation for Robust Network by Adversarial Instrumental Variable Regression

CVPR 2023poster

The origin of adversarial examples is still inexplicable in research fields, and it arouses arguments from various viewpoints, albeit comprehensive investigations. In this paper, we propose a way of delving into the unexpected vulnerability in adversarially trained networks from a causal perspective…

2023

Diffusion Video Autoencoders: Toward Temporally Consistent Face Video Editing via Disentangled Video Encoding

CVPR 2023poster

Inspired by the impressive performance of recent face image editing methods, several studies have been naturally proposed to extend these methods to the face video editing task. One of the main challenges here is temporal consistency among edited frames, which is still unresolved. To this end, we pr…

Cited by 34SourcePDFScholar
2023

Dynamic Structure Pruning for Compressing CNNs

AAAI 2023technical

Structure pruning is an effective method to compress and accelerate neural networks. While filter and channel pruning are preferable to other structure pruning methods in terms of realistic acceleration and hardware compatibility, pruning methods with a finer granularity, such as intra-channel pruni…

2023

Leap-of-Thought: Accelerating Transformers via Dynamic Token Routing

EMNLP 2023long main

Computational inefficiency in transformers has been a long-standing challenge, hindering the deployment in resource-constrained or real-time applications. One promising approach to mitigate this limitation is to progressively remove less significant tokens, given that the sequence length strongly co…

Cited by 0SourceScholar
2023

Learning Input-agnostic Manipulation Directions in StyleGAN with Text Guidance

ICLR 2023poster

With the advantages of fast inference and human-friendly flexible manipulation, image-agnostic style manipulation via text guidance enables new applications that were not previously available. The state-of-the-art text-guided image-agnostic manipulation method embeds the representation of each chann…

2023

Mitigating Adversarial Vulnerability through Causal Parameter Estimation by Adversarial Double Machine Learning

ICCV 2023poster

Adversarial examples derived from deliberately crafted perturbations on visual inputs can easily harm decision process of deep neural networks. To prevent potential threats, various adversarial training-based defense methods have grown rapidly and become a de facto standard approach for robustness.…

Cited by 11PDFcodeScholar
2023

Rarity Score : A New Metric to Evaluate the Uncommonness of Synthesized Images

ICLR 2023top-25%

Evaluation metrics in image synthesis play a key role to measure performances of generative models. However, most metrics mainly focus on image fidelity. Existing diversity metrics are derived by comparing distributions, and thus they cannot quantify the diversity or rarity degree of each generated…

Cited by 36SourcePDFScholar
2023

SMoP: Towards Efficient and Effective Prompt Tuning with Sparse Mixture-of-Prompts

EMNLP 2023short main

Prompt tuning has emerged as a successful parameter-efficient alternative to the full fine-tuning of language models. However, prior works on prompt tuning often utilize long soft prompts of up to 100 tokens to improve performance, overlooking the inefficiency associated with extended inputs. In thi…

Cited by 28SourcecodeScholar
2022

CPO: Change Robust Panorama to Point Cloud Localization

ECCV 2022poster

"We present CPO, a fast and robust algorithm that localizes a 2D panorama with respect to a 3D point cloud of a scene possibly containing changes. To robustly handle scene changes, our approach deviates from conventional feature point matching, and focuses on the spatial context provided from panora…

2022

Efficient Pre-training of Masked Language Model via Concept-based Curriculum Masking

EMNLP 2022main

Self-supervised pre-training has achieved remarkable success in extensive natural language processing tasks. Masked language modeling (MLM) has been widely used for pre-training effective bidirectional representations but comes at a substantial training cost. In this paper, we propose a novel concep…

2022

Generating Videos with Dynamics-aware Implicit Generative Adversarial Networks

ICLR 2022poster

In the deep learning era, long video generation of high-quality still remains challenging due to the spatio-temporal complexity and continuity of videos. Existing prior works have attempted to model video distribution by representing videos as 3D grids of RGB values, which impedes the scale of gener…

Cited by 234SourcePDFScholar
2022

Generator Knows What Discriminator Should Learn in Unconditional GANs

ECCV 2022poster

"Recent methods for conditional image generation benefit from dense supervision such as segmentation label maps to achieve high-fidelity. However, it is rarely explored to employ dense supervision for unconditional image generation. Here we explore the efficacy of dense supervision in unconditional…

2022

Learning from Missing Relations: Contrastive Learning with Commonsense Knowledge Graphs for Commonsense Inference

ACL 2022findings

Commonsense inference poses a unique challenge to reason and generate the physical, social, and causal conditions of a given event. Existing approaches to commonsense inference utilize commonsense transformers, which are large-scale language models that learn commonsense knowledge graphs. However, t…

2022

Masking Adversarial Damage: Finding Adversarial Saliency for Robust and Sparse Network

CVPR 2022poster

Adversarial examples provoke weak reliability and potential security issues in deep neural networks. Although adversarial training has been widely studied to improve adversarial robustness, it works in an over-parameterized regime and requires high computations and large memory budgets. To bridge ad…

Cited by 18PDFcodeScholar
2022

MoDA: Map Style Transfer for Self-Supervised Domain Adaptation of Embodied Agents

ECCV 2022poster

"We propose a domain adaptation method, MoDA, which adapts a pretrained embodied agent to a new, noisy environment without ground-truth supervision. Map-based memory provides important contextual information for visual navigation, and exhibits unique spatial structure mainly composed of flat walls a…

Cited by 10SourcePDFScholar
2022

Tutoring Helps Students Learn Better: Improving Knowledge Distillation for BERT with Tutor Network

EMNLP 2022main

Pre-trained language models have achieved remarkable successes in natural language processing tasks, coming at the cost of increasing model size. To address this issue, knowledge distillation (KD) has been widely applied to compress language models. However, typical KD approaches for language models…

Cited by 5SourcePDFScholar
2021

CTRL-C: Camera Calibration TRansformer With Line-Classification

ICCV 2021poster

Single image camera calibration is the task of estimating the camera parameters from a single input image, such as the vanishing points, focal length, and horizon line. In this work, we propose Camera calibration TRansformer with Line-Classification (CTRL-C), an end-to-end neural network-based appro…

Cited by 49PDFcodeScholar
2021

Distilling Robust and Non-Robust Features in Adversarial Examples by Information Bottleneck

NeurIPS 2021poster

Adversarial examples, generated by carefully crafted perturbation, have attracted considerable attention in research fields. Recent works have argued that the existence of the robust and non-robust features is a primary cause of the adversarial examples, and investigated their internal interactions…

2021

Exploiting Spatial Dimensions of Latent in GAN for Real-Time Image Editing

CVPR 2021poster

Generative adversarial networks (GANs) synthesize realistic images from random latent vectors. Although manipulating the latent vectors controls the synthesized outputs, editing real images with GANs suffers from i) time-consuming optimization for projecting real images to the latent vectors, ii) or…

Cited by 192PDFcodeScholar
2021

N-ImageNet: Towards Robust, Fine-Grained Object Recognition With Event Cameras

ICCV 2021poster

We introduce N-ImageNet, a large-scale dataset targeted for robust, fine-grained object recognition with event cameras. The dataset is collected using programmable hardware in which an event camera consistently moves around a monitor displaying images from ImageNet. N-ImageNet serves as a challengin…

Cited by 105PDFcodeScholar
2020

Neural Geometric Parser for Single Image Camera Calibration

ECCV 2020poster

We propose a neural geometric parser learning single image camera calibration for man-made scenes. Unlike previous neural approaches that rely only on semantic cues obtained from neural networks, our approach considers both semantic and geometric cues, resulting in significant accuracy improvement.…

2020

U-GAT-IT: Unsupervised Generative Attentional Networks with Adaptive Layer-Instance Normalization for Image-to-Image Translation

ICLR 2020poster

We propose a novel method for unsupervised image-to-image translation, which incorporates a new attention module and a new learnable normalization function in an end-to-end manner. The attention module guides our model to focus on more important regions distinguishing between source and target domai…

Cited by 764SourceScholar